arXiv:2503.15627eess.ASeess.SP2025-03

建立雷达测振与语音声学的数学关联,提升雷达语音感知性能。

A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations

  • 构建雷达测得颈部振动与语音声学的解析关系模型。
  • 66人实验显示雷达信号更贴近模型滤波后的振动信号。
  • 适用于语音增强、安全认证等需要非接触语音感知场景。

毫米波(mmWave)雷达已成为语音感知的有前景模态,相较于传统麦克风具有优势。以往研究证实雷达可捕捉与发声相关的运动信号,但雷达测得的振动与语音声学之间的分析关联尚不清晰。本文建立了一个数学框架,将雷达捕获的颈部振动与语音声学相联系,推导出颈部表面位移与语音之间的解析关系。基于66名受试者的数据,通过统计谱距离分析对模型进行实证评估。结果表明,雷达测得信号与由语音生成的模型滤波振动信号的匹配度,高于与原始语音本身的匹配度。该成果为雷达语音处理在语音增强、编码、监控及身份认证等应用中的改进提供了理论基础。

原文摘要 · Abstract (English)

Millimeter Wave (mmWave) radar has emerged as a promising modality for speech sensing, offering advantages over traditional microphones. Prior works have demonstrated that radar captures motion signals related to vocal vibrations, but there is a gap in the understanding of the analytical connection between radar-measured vibrations and acoustic speech signals. We establish a mathematical framework linking radar-captured neck vibrations to speech acoustics. We derive an analytical relationship between neck surface displacements and speech. We use data from 66 human participants, and statistical spectral distance analysis to empirically assess the model. Our results show that the radar-measured signal aligns more closely with our model filtered vibration signal derived from speech than with raw speech itself. These findings provide a foundation for improved radar-based speech processing for applications in speech enhancement, coding, surveillance, and authentication.

雷达语音非接触感知语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。