arXiv:2502.05762eess.AS2025-02ACL被引 3

无需音频依赖,直接将无声口动肌电信号转为文本。

Non-invasive electromyographic speech neuroprosthesis: a geometric perspective

  • 用多部位表面肌电捕捉无声口动信号,构建高效高维表示。
  • 首次实现无音频对齐的肌电到语音级文本端到端转换。
  • 适合失语或声带切除患者,有望恢复自然沟通能力。

我们提出一种神经肌肉言语接口,可将无声发声动作直接转化为文本。通过在面部和颈部多个发音部位采集表面肌电(EMG)信号,记录受试者无声口动时的肌电信号,实现直接的肌电到文本转换。该接口有望帮助因喉切除、神经肌肉疾病、中风或创伤性损伤(如放疗毒性)导致发音器官受损而丧失清晰言语能力的人群恢复交流能力。以往研究多聚焦于将可听发声时的肌电信号映射到时间对齐的音频目标,或将其转移至无声肌电记录,但此类方法依赖音频,无法应用于完全失语患者。相比之下,本文提出一种高效的高维肌电信号表示方法,并展示了无需时间对齐音频即可实现的语音级端到端肌电到文本转换。

原文摘要 · Abstract (English)

We present a neuromuscular speech interface that translates silently voiced articulations directly into text. We record surface electromyographic (EMG) signals from multiple articulatory sites on the face and neck as participants silently articulate speech, enabling direct EMG-to-text translation. Such an interface has the potential to restore communication for individuals who have lost the ability to produce intelligible speech due to laryngectomy, neuromuscular disease, stroke, or trauma-induced damage (e.g., radiotherapy toxicity) to the speech articulators. Prior work has largely focused on mapping EMG collected during audible articulation to time-aligned audio targets or transferring these targets to silent EMG recordings, which inherently requires audio and limits applicability to patients who can no longer speak. In contrast, we propose an efficient representation of high-dimensional EMG signals and demonstrate direct sequence-to-sequence EMG-to-text conversion at the phonemic level without relying on time-aligned audio.

肌电接口无声言语神经假体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。