通过表面肌电捕捉情绪,无声说话时也能识别沮丧情绪。
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
- 用面部和颈部肌电信号分析情绪表达,跨说话模式有效
- 识别沮丧情绪最高达0.845 AUC,无声时仍可准确解码
- 适合开发不依赖声音的情绪感知语音接口
情感表达是口语交流的核心,但其与发音动作之间的关联尚不明确。表面肌电(sEMG)等肌肉活动测量方法可能揭示情绪如何调节言语产生,同时结合声学分析。本文研究在发声与无声言语状态下,从面部和颈部的表面肌电(sEMG)中解码情感。为此,我们构建了一个包含12名参与者、3项任务共2,780条语句的数据集,并采用多种特征与模型嵌入进行个体内与个体间解码评估。结果表明,肌电表征能可靠地区分沮丧情绪,最大AUC达0.845,且在不同发音模式间具有良好泛化能力。消融实验进一步证明,情感特征存在于面部运动活动中,并在无发音时依然存在,凸显了肌电传感在情绪感知型无声语音接口中的潜力。
原文摘要 · Abstract (English)
The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech production. For this purpose, we introduce a dataset comprising 2,780 utterances from 12 participants across 3 tasks, on which we evaluate both intra- and inter-subject decoding using a range of features and model embeddings. Our results reveal that EMG representations reliably discriminate frustration with up to 0.845 AUC, and generalize well across articulation modes. Our ablation study further demonstrates that affective signatures are embedded in facial motor activity and persist in the absence of phonation, highlighting the potential of EMG sensing for affect-aware silent speech interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。