arXiv:2607.06014cs.SD2026-07

让语音模型更懂说话人差异,提升多轮问答准确率

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

论文配图:Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
图 1 · 摘自论文原文
  • 将查询分组并约束方向,防止特征向量坍缩
  • 在SAKURA数据集上准确率提升至75.2%,超基线26.4点
  • 显著增强不同说话人的区分度,适合语音对话系统

音频-语言模型通过查询变换器(Q-Former)压缩语音编码器输出后输入大语言模型。我们发现该压缩过程存在两个问题:输出向量趋于单一方向,且不同说话人产生几乎无法区分的表示,声调、性别等副语言特征丢失。提出ORCA方法,将查询分组并约束其输出指向不同方向,以逆转特征坍缩。在SAKURA多跳推理任务中,ORCA相比相同训练的4B基线提升26.4个百分点,达75.2%(8B Audio Flamingo-3为49.0%)。在连接器层面,该改进使查询冗余降低12倍,跨说话人方差提升75倍。

原文摘要 · Abstract (English)

Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures in this compression. The connector's output vectors collapse to a single direction, and different speakers produce nearly indistinguishable outputs, with paralinguistic cues such as speaker identity, gender, and prosody lost along the way. Our method, ORCA, reverses this collapse by splitting the queries into groups whose outputs are constrained to point in different directions. On SAKURA multi-hop reasoning, ORCA gains 26.4 points over an identically trained 4B baseline, reaching 75.2% (vs. 49.0% for the 8B Audio Flamingo-3). At the connector level, the same change cuts query redundancy by 12x and raises cross-speaker variance by 75x.

音频理解多模态语音表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。