arXiv:2603.11877eess.AS2026-03综述

用大脑肌肉信号实现无声对话,让设备读懂你的心思。

Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review

  • 用神经、肌肉、口动等信号直接解码语言意图,跳过说话环节。
  • 结合大模型后,错误率接近实用门槛,可部署在耳机眼镜等穿戴设备上。
  • 解决个体差异难题,推动无声接口向隐私安全的智能未来演进。

人机交互长期依赖声学通道,易受噪声干扰、侵犯隐私且不适用于言语障碍者。无声语音接口(SSIs)通过直接从神经-肌肉-发音连续体中解码语言意图,突破这一局限。本文系统梳理了从传统传感器中心分析到整体‘意图到执行’分类框架的演进,评估了四类生理截获点:神经振荡、神经肌肉激活、发音运动学(超声/磁测)以及基于声波或射频的主动探测。关键转折在于从启发式信号处理转向潜在语义对齐,利用大语言模型(LLMs)和生成架构作为高层语义先验,缓解生物信号的信息稀疏与非平稳问题。现代框架首次将零散的生理动作映射至结构化语义潜空间,使词错误率(WER)达到实际部署可用水平。此外,研究还探讨了从实验室设备向耳戴式、智能眼镜等消费级可穿戴设备的隐形接口转型。最后,提出通过自监督基础模型应对‘用户依赖悖论’,并界定‘神经安全’伦理边界,以保障认知自由。

原文摘要 · Abstract (English)

Human-computer interaction has traditionally relied on the acoustic channel, a dependency that introduces systemic vulnerabilities to environmental noise, privacy constraints, and physiological speech impairments. Silent Speech Interfaces (SSIs) emerge as a transformative paradigm that bypasses the acoustic stage by decoding linguistic intent directly from the neuro-muscular-articulatory continuum. This review provides a high-level synthesis of the SSI landscape, transitioning from traditional transducer-centric analysis to a holistic intent-to-execution taxonomy. We systematically evaluate sensing modalities across four critical physiological interception points: neural oscillations, neuromuscular activation, articulatory kinematics (ultrasound/magnetometry), and pervasive active probing via acoustic or radio-frequency sensing. Critically, we analyze the current paradigm shift from heuristic signal processing to Latent Semantic Alignment. In this new era, Large Language Models (LLMs) and deep generative architectures serve as high-level linguistic priors to resolve the ``informational sparsity'' and non-stationarity of biosignals. By mapping fragmented physiological gestures into structured semantic latent spaces, modern SSI frameworks have, for the first time, approached the Word Error Rate usability threshold required for real-world deployment. We further examine the transition of SSIs from bulky laboratory instrumentation to ``invisible interfaces'' integrated into commodity-grade wearables, such as earables and smart glasses. Finally, we outline a strategic roadmap addressing the ``user-dependency paradox'' through self-supervised foundation models and define the ethical boundaries of ``neuro-security'' to protect cognitive liberty in an increasingly interfaced world.

无声语音脑机接口大模型可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。