arXiv:2609.02165eess.AS2026-09

用耳塞振动捕捉自言自语,提升嘈杂环境下的语音清晰度

Sensing Bone-Conducted Speech with Earbuds

论文配图:Sensing Bone-Conducted Speech with Earbuds
图 1 · 摘自论文原文
  • 通过测量耳塞壳体振动,分析自言自语引起的骨传导信号特性
  • 振动频谱呈低通特性,400 Hz以上衰减达93 dB/十倍频程
  • 单轴传感器在400 Hz以下可实现低于1.5 dB的信号损失

在使用耳塞进行移动通信时,清晰捕捉佩戴者自言自语(OV)至关重要,但在嘈杂环境中仍具挑战性。利用耳塞外壳振动感知骨传导(BC)语音可改善OV捕获效果。然而,此前尚未详细分析OV引起耳塞振动的带宽与空间特性,而这些特性对传感器选型与位置布局至关重要。本研究基于两种耳塞模型的实测数据,揭示了其振动特性:频域上,振动呈现低通特征,400 Hz以上衰减高达-93 dB/十倍频程,因此需采用噪声底噪较低的传感器以感知1 kHz以上的振动;空间上,耳塞主要沿耳道入口方向前后振动,跨受试者和实验组间具有高度一致性。仿真表明,单轴传感器即可在400 Hz以下实现小于1.5 dB的平均信号衰减。

原文摘要 · Abstract (English)

Clear capture of the wearer's own voice (OV) is essential when using earbuds for mobile communication. However, OV capture remains challenging in noisy environments. Bone-conducted (BC) speech, which can be sensed as vibrations of the earbud housing, can be used to improve OV capture. However, neither bandwidth nor spatial characteristics of OV-induced earbud vibrations have been analyzed in detail, despite both characteristics being relevant, e.g., for sensor choice and placement. This study investigates both characteristics, based on measurements with two earbud models. Spectrally, results indicate that OV-induced earbud vibrations exhibit a low-pass characteristic, with a steep roll-off of -93 dB per decade above 400 Hz. Thus, sensors with comparatively low noise floors are required to sense the vibrations above \SI{1}{\kilo\hertz}. Spatially, results indicate that the earbuds mainly vibrate in and out of the ear canal entrance, with high consistency between subjects and fits. Simulations confirm that this enables capture of the high-power vibrations below 400 Hz by a single-axis sensor with less than 1.5 dB mean attenuation.

骨传导语音增强耳塞传感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。