融合Whisper与Qwen,动态选择音频语义特征
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation

- 用门控注意力机制动态融合双模型特征
- 在跨域音频任务上达到0.836的最优得分
- 适合需要多模态音频理解的场景
尽管预训练模型在特定任务上表现优异,但在跨不同声学领域学习通用表示仍具挑战。为此,我们提出WQ-Fusion,一种稳健的双编码器框架,用于跨域音频表征学习。克服静态拼接的局限,WQ-Fusion通过自适应特征调制模块和新颖的逐元素门控注意力机制,融合Whisper与Qwen。该设计实现动态特征选择,使模型可有侧重地强调相关声学与语义维度。在Interspeech 2026音频编码器能力挑战赛(第A赛道)基准上的大量实验表明,通过有效路由异构信息,WQ-Fusion取得了0.836的优异综合得分,显著优于最强单编码器基线。
原文摘要 · Abstract (English)
While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion, a robust dual-encoder framework for cross-domain audio representation learning. Overcoming the limitations of static concatenation, WQ-Fusion integrates whisper and qwen via an Adaptive Feature Modulation module and a novel element-wise gated attention mechanism. This design enables dynamic feature selection, allowing the model to selectively emphasize relevant acoustic and semantic dimensions. Extensive experiments on the Interspeech 2026 Audio Encoder Capability Challenge (Track A) benchmark demonstrate that by effectively routing heterogeneous information, WQ-Fusion achieves a superior overall score of 0.836, significantly outperforming the strongest single-encoder baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。