arXiv:2603.10324cs.HCcs.AI2026-03被引 1

鼻戴式麦克风+振动传感器,实现安静语音交互

NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction

  • 鼻梁位置融合声学与振动信号,捕捉低声语
  • 在噪声环境下仍保持高识别率和语音质量
  • 适合需要持续、隐蔽语音交互的用户

静默与低语语音为始终可用的语音交互带来可能,但现有方法难以兼顾词汇量、佩戴舒适性、静音效果与抗噪能力。我们提出NasoVoce,一种佩戴于智能眼镜鼻托处的接口,集成麦克风与振动传感器。该位置靠近口部,可无感获取骨传导与皮肤传导的语音信号,可靠捕获低音量语音(如耳语)。麦克风虽能获取高质量音频,但易受环境噪声干扰;振动传感器抗噪性强,但信号质量较低。通过融合互补信号,NasoVoce生成高鲁棒性语音。使用Whisper Large-v2、PESQ、STOI和MUSHRA评估显示,其语音识别准确率与音质显著提升。实验证明,该系统具备实现持续、隐蔽、始终可用的AI语音对话的可行性。

原文摘要 · Abstract (English)

Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-bridge-mounted interface that integrates a microphone and a vibration sensor. Positioned at the nasal pads of smart glasses, it unobtrusively captures both acoustic and vibration signals. The nasal bridge, close to the mouth, allows access to bone- and skin-conducted speech and enables reliable capture of low-volume utterances such as whispered speech. While the microphone captures high-quality audio, it is highly sensitive to environmental noise. Conversely, the vibration sensor is robust to noise but yields lower signal quality. By fusing these complementary inputs, NasoVoce generates high-quality speech robust against interference. Evaluation with Whisper Large-v2, PESQ, STOI, and MUSHRA ratings confirms improved recognition and quality. NasoVoce demonstrates the feasibility of a practical interface for always-available, continuous, and discreet AI voice conversations.

语音交互可穿戴设备低噪音耳语识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。