arXiv:2508.19483eess.AS2025-08被引 3

用视听同步提升助听器语音增强效果,让嘈杂环境下的对话更清晰。

Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids

  • 轻量级跨注意力模型实现音视频特征动态对齐。
  • 实时处理仅36毫秒延迟,PESQ达0.52,STOI提升19%。
  • 适合需要低延迟、高语音可懂度的助听设备应用。

针对助听器中实时语音增强的挑战,本文提出一种视听特征同步方法,通过融合听觉信号与视觉线索,利用双模态互补性提升语音可懂度,尤其在强噪声环境中表现突出。研究设计了一种轻量级跨注意力模型,基于大规模数据和简洁架构学习鲁棒的音视频联合表示。该模型嵌入到视听语音增强(AVSE)框架中,能动态聚焦关键音视频特征,实现精准同步,显著提升语音质量。实验在AVSEC3数据集上验证,系统在保持极低延迟(36ms)和能耗的同时,实现显著性能提升:感知质量(PESQ)达0.52,语音可懂度(STOI)提高19%,保真度(SI-SDR)达10.10dB,优于现有基线方法。

原文摘要 · Abstract (English)

Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach integrates auditory signals with visual cues, utilizing the complementary description of these modalities to improve speech intelligibility. Audio-visual feature synchronization for real-time SE in hearing aids can be further optimized using an efficient feature alignment module. In this study, a lightweight cross-attentional model learns robust audio-visual representations by exploiting large-scale data and simple architecture. By incorporating the lightweight cross-attentional model in an AVSE framework, the neural system dynamically emphasizes critical features across audio and visual modalities, enabling defined synchronization and improved speech intelligibility. The proposed AVSE model not only ensures high performance in noise suppression and feature alignment but also achieves real-time processing with minimal latency (36ms) and energy consumption. Evaluations on the AVSEC3 dataset show the efficiency of the model, achieving significant gains over baselines in perceptual quality (PESQ:0.52), intelligibility (STOI:19\%), and fidelity (SI-SDR:10.10dB).

视听融合语音增强助听器低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。