用仿真声学模型让单麦克风助听器精准识别用户自己的声音
Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids
- 通过仿真声学传递函数数据增强,训练单麦克风语音检测模型
- 在真实助听器数据上达80%准确率,一秒钟语音仍保持90%精度
- 无需复杂硬件,适合低成本智能助听器研发
本文提出一种基于仿真的单麦克风自声检测(OVD)方法,用于提升助听器用户体验。现有方案多依赖多麦克风或额外传感器,增加设备复杂性与成本。为实现无需昂贵传输函数测量的机器学习型OVD,研究设计了一种基于仿真声学传递函数(ATFs)的数据增强策略,使模型覆盖多种空间传播条件。采用基于Transformer的分类器,先在解析生成的ATFs上训练,再逐步微调至数值仿真的人体头身模型。该分层适应机制有效提升模型的空间理解能力并保持泛化性能。实验表明,在人体头身模型测试数据上准确率达95.52%;短时语音(1秒)下仍保持90.02%准确率;在真实助听器录音上未微调即达80%准确率,得益于轻量级测试时特征补偿。结果验证了从仿真到现实的强泛化能力,展示了该方法在助听器设计中的实用前景。
原文摘要 · Abstract (English)
This paper presents a simulation-based approach to own voice detection (OVD) in hearing aids using a single microphone. While OVD can significantly improve user comfort and speech intelligibility, existing solutions often rely on multiple microphones or additional sensors, increasing device complexity and cost. To enable ML-based OVD without requiring costly transfer-function measurements, we propose a data augmentation strategy based on simulated acoustic transfer functions (ATFs) that expose the model to a wide range of spatial propagation conditions. A transformer-based classifier is first trained on analytically generated ATFs and then progressively fine-tuned using numerically simulated ATFs, transitioning from a rigid-sphere model to a detailed head-and-torso representation. This hierarchical adaptation enabled the model to refine its spatial understanding while maintaining generalization. Experimental results show 95.52% accuracy on simulated head-and-torso test data. Under short-duration conditions, the model maintained 90.02% accuracy with one-second utterances. On real hearing aid recordings, the model achieved 80% accuracy without fine-tuning, aided by lightweight test-time feature compensation. This highlights the model's ability to generalize from simulated to real-world conditions, demonstrating practical viability and pointing toward a promising direction for future hearing aid design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。