arXiv:2509.20875eess.AS2025-09中稿 · ICASSP 2026

融合个人语音与体感麦克风,提升耳机语音降噪效果

PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables

  • 用用户录音和耳内麦克风双路输入增强语音
  • 双路结合使降噪性能显著优于单一方法
  • 即使录入时有噪音,仍比纯体感方案更优

可穿戴设备中的语音拾取需在抑制噪声和干扰说话人同时保持用户自身语音质量。单通道方法难以区分目标语音与干扰源,本文对比两种策略:个性化语音增强(PSE)利用用户注册语音作为目标表征,辅助传感器语音增强(AS-SE)则使用耳内麦克风作为额外输入。在两个公开数据集上评估不同辅助传感器阵列的跨数据集泛化能力,并提出训练时增强策略以提升AS-SE系统泛化性。结果表明,融合PSE与AS-SE(PAS-SE)能带来互补优势,尤其当注册语音由耳内麦克风录制时。进一步证明,使用带噪耳内录音进行个性化训练的PAS-SE,性能仍优于纯AS-SE系统。

原文摘要 · Abstract (English)

Speech enhancement for voice pickup in hearables aims to improve the user's voice by suppressing noise and interfering talkers, while maintaining own-voice quality. For single-channel methods, it is particularly challenging to distinguish the target from interfering talkers without additional context. In this paper, we compare two strategies to resolve this ambiguity: personalized speech enhancement (PSE), which uses enrollment utterances to represent the target, and auxiliary-sensor speech enhancement (AS-SE), which uses in-ear microphones as additional input. We evaluate the strategies on two public datasets, employing different auxiliary sensor arrays, to investigate their cross-dataset generalization. We propose training-time augmentations to facilitate cross-dataset generalization of AS-SE systems. We also show that combining PSE and AS-SE (PAS-SE) provides complementary performance benefits, especially when enrollment speech is recorded with the in-ear microphone. We further demonstrate that PAS-SE personalized with noisy in-ear enrollments maintains performance benefits over the AS-SE system.

语音增强可穿戴设备多模态个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。