arXiv:2602.06460cs.SD2026-02

减少肌电通道数仍能高效还原语音,关键在通道组合与预训练设计。

EMG-to-Speech with Fewer Channels

  • 通过分析通道互补性,找到最优少通道组合提升语音重建效果。
  • 4-6通道设置下微调性能优于从零训练,最高提升12.7%(相对)。
  • 适合开发轻量化、实用化的无声语音接口系统。

表面肌电图(EMG)是无声语音接口的有前景技术,但其效果受传感器位置和通道数量影响显著。本文研究单个及组合EMG通道对语音重建的贡献,发现某些通道单独更有效,而最佳性能来自具有互补性的通道子集。通过通道消融实验分析音素分类准确率,观察到反映肌肉解剖功能的可解释模式。为缓解通道减少导致的性能下降,采用随机通道丢弃策略在8通道全数据上预训练模型,并在减少通道的子集上微调。微调在4-6通道设置中持续优于从零训练,最优丢弃策略随通道数变化。结果表明,通过预训练与通道感知设计可缓解传感器缩减带来的性能损失,支持轻量化、实用化EMG无声语音系统的开发。

原文摘要 · Abstract (English)

Surface electromyography (EMG) is a promising modality for silent speech interfaces, but its effectiveness depends heavily on sensor placement and channel availability. In this work, we investigate the contribution of individual and combined EMG channels to speech reconstruction performance. Our findings reveal that while certain EMG channels are individually more informative, the highest performance arises from subsets that leverage complementary relationships among channels. We also analyzed phoneme classification accuracy under channel ablations and observed interpretable patterns reflecting the anatomical roles of the underlying muscles. To address performance degradation from channel reduction, we pretrained models on full 8-channel data using random channel dropout and fine-tuned them on reduced-channel subsets. Fine-tuning consistently outperformed training from scratch for 4 - 6 channel settings, with the best dropout strategy depending on the number of channels. These results suggest that performance degradation from sensor reduction can be mitigated through pretraining and channel-aware design, supporting the development of lightweight and practical EMG-based silent speech systems.

肌电语音少通道预训练无声接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。