arXiv:2604.18920cs.SDcs.CL2026-04

SPARC特征比音素特征更准预测肌电信号,跨发音模式表现稳定。

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

论文配图:Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features
图 1 · 摘自论文原文
  • 用SPARC和音素特征线性建模肌电信号,跨模式验证
  • SPARC在所有电极和模式下均优于音素表示,子声言语仍显著高于随机
  • 结果具解剖可解释性,适合静默语音建模研究者参考

我们测试了语音构音编码(SPARC)特征是否能在线性模型中准确预测24名受试者在出声、默念和子声三种发音模式下的表面肌电(sEMG)包络。采用弹性网多变量时间响应函数(mTRF)进行句子级交叉验证,结果显示SPARC在几乎所有电极和所有发音模式下均优于音素独热编码。出声与默念表现相近,子声仍显著高于随机水平,表明其存在可检测的构音活动。方差分解显示SPARC有显著独特贡献,而音素特征贡献极小。mTRF权重模式揭示了电极位置与构音运动之间的解剖可解释关系,且在不同模式间保持一致。本研究聚焦表征/编码分析(非端到端解码),支持SPARC作为基于sEMG的静默语音建模的稳健且可解释的中间目标。

原文摘要 · Abstract (English)

We test whether Speech Articulatory Coding (SPARC) features can linearly predict surface electromyography (sEMG) envelopes across aloud, mimed, and subvocal speech in twenty-four subjects. Using elastic-net multivariate temporal response function (mTRF) with sentence-level cross-validation, SPARC yields higher prediction accuracy than phoneme one-hot representations on nearly all electrodes and in all speech modes. Aloud and mimed speech perform comparably, and subvocal speech remains above chance, indicating detectable articulatory activity. Variance partitioning shows a substantial unique contribution from SPARC and a minimal unique contribution from phoneme features. mTRF weight patterns reveal anatomically interpretable relationships between electrode sites and articulatory movements that remain consistent across modes. This study focuses on representation/encoding analysis (not end-to-end decoding) and supports SPARC as a robust and interpretable intermediate target for sEMG-based silent-speech modeling.

肌电预测语音编码静默语音构音特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。