arXiv:2506.01148eess.AScs.SD2025-06中稿 · INTERSPEECH 2025被引 1

用强化学习注意力融合音频编码器与频谱特征,提升心脏杂音分类性能。

Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism

  • 设计基于强化学习的跨注意力机制,动态选择关键特征头
  • 在公开数据集上达到新最高准确率,显著优于单一特征或传统融合方法
  • 适合医疗听诊分析、多模态信号处理方向的研究者参考

本研究聚焦心脏杂音分类(HMC),假设将神经音频编码器表示(NACRs,如EnCodec)与频谱特征(SFs,如MFCC)融合可带来更优性能。我们认为二者具有互补性:NACRs擅长捕捉节奏变化等精细声学模式,而SFs则关注谐波结构、频域能量分布等关键频率特性。为此,我们提出BAOMI框架,采用新型基于强化学习的跨注意力机制实现高效融合。该机制通过智能代理为多头注意力中最重要的头分配更高权重,有效抑制噪声干扰。实验表明,BAOMI在多个基准测试中均取得最佳表现,超越单独使用NACRs、SFs或传统融合方法,刷新了当前最优性能记录。

原文摘要 · Abstract (English)

In this study, we focus on heart murmur classification (HMC) and hypothesize that combining neural audio codec representations (NACRs) such as EnCodec with spectral features (SFs), such as MFCC, will yield superior performance. We believe such fusion will trigger their complementary behavior as NACRs excel at capturing fine-grained acoustic patterns such as rhythm changes, spectral features focus on frequency-domain properties such as harmonic structure, spectral energy distribution crucial for analyzing the complex of heart sounds. To this end, we propose, BAOMI, a novel framework banking on novel bandit-based cross-attention mechanism for effective fusion. Here, a agent provides more weightage to most important heads in multi-head cross-attention mechanism and helps in mitigating the noise. With BAOMI, we report the topmost performance in comparison to individual NACRs, SFs, and baseline fusion techniques and setting new state-of-the-art.

心音分类特征融合注意力机制医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。