通过掩码融合肌电与口型信息,提升无声语音合成的鲁棒性。
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

- 训练时对肌电和口型信号进行模态掩码,实现联合建模。
- 在多说话人场景下,词错误率降低14个百分点。
- 对传感器失效或信号退化有更强适应性,适合真实应用。
通过无声语音接口(SSIs)恢复语音已成为帮助声带功能受损或缺失者的重要辅助技术。非侵入式模态中,表面肌电图(sEMG)和基于视频的口型识别可提供互补的发音信息,但两者在连续语音合成中的融合仍研究不足。现有多模态方法很少考虑模态退化或临时传感器故障下的鲁棒性,限制了实际应用。本文提出一种掩码多模态语音合成框架,通过训练时对sEMG和口型信号进行模态掩码,实现联合利用。在多说话人设置下,该方法相比最强单模态基线,词错误率降低最多达14个百分点。实验表明,掩码策略对性能提升和低比特率条件下的鲁棒性至关重要,且在模态缺失情况下优于特定退化数据增强。音素级分析进一步揭示了两模态的互补性,尤其对元音及特定辅音组收益显著。总体证明掩码多模态融合在无声语音合成中有效且鲁棒,但针对喉切除患者适配仍是开放挑战。
原文摘要 · Abstract (English)
Speech restoration through silent speech interfaces (SSIs) has emerged as a promising assistive technology for individuals with impaired or absent laryngeal voice production. Among non-invasive SSI modalities, surface electromyography (sEMG) and video-based lipreading provide complementary articulatory information, yet their integration for continuous speech synthesis remains underexplored. Moreover, existing multimodal approaches rarely address robustness to modality degradation or temporary sensor failure, limiting their applicability in realistic scenarios. In this work, we propose a masked multimodal speech synthesis framework that jointly leverages sEMG and lipreading signals through modality masking during training. Under multispeaker settings, the proposed approach reduces word error rate by up to 14 absolute percentage points compared to the strongest unimodal baseline. Experimental results not only show that masking strategies are critical for these performance gains and robustness under low-bitrate conditions, but also that they generalize better than degradation-specific data augmentations in the presence of modality absence conditions. Phone-level analyses further reveal complementary contributions across modalities, with particularly strong benefits for vowels and for specific consonant groups. Overall, these findings demonstrate the effectiveness and robustness of masked multimodal integration for silent speech synthesis, although adaptation to laryngectomized speakers remains an open research challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。