用强化学习对齐语音与文本推理路径,缩小语音大模型的推理差距。
Closing the Modality Reasoning Gap for Speech Large Language Models
- 设计非对称奖励机制,对齐语音与文本输入的推理轨迹。
- 在MMSU和OBQA上显著提升语音推理表现,达70亿级语音大模型最优水平。
- 适合关注多模态大模型推理能力优化的研究者与开发者。
尽管语音大语言模型已取得显著进展,但其在语音输入上的推理能力仍明显弱于文本输入,存在显著的模态推理差距。该差距可能源于Transformer层间的表征漂移以及长链推理中的行为偏差。为此,本文提出TARS,一种通过非对称奖励设计对齐文本条件与语音条件轨迹的强化学习框架。该框架引入两个密集且互补的信号:表征对齐,衡量语音与文本条件轨迹在各层隐藏状态间的相似性;行为对齐,评估生成输出与参考文本补全之间的语义一致性。在MMSU和OBQA等高难度推理基准上的实验表明,该方法显著缩小了模态推理差距,并在70亿规模的语音大模型中达到当前最佳性能。
原文摘要 · Abstract (English)
Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。