用文本推理知识蒸馏提升音频模型逻辑能力
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
- 让文本教师指导音频学生生成带推理路径的回应
- 在多个数据集上显著提升音频推理准确率和思维链质量
- 适合研究多模态推理与语音理解的学者
尽管大型音频语言模型在听觉感知方面取得了显著进展,但在深层逻辑推理方面仍远落后于文本型大模型,主要由于高质量音频推理数据稀缺。为此,我们提出 X³-OPD,一种跨模态的在线对齐知识蒸馏框架,将强大文本教师的推理能力迁移至音频语言学生模型。训练过程中,学生基于自身声学感知生成推理轨迹,教师则利用匹配的文本输入和验证答案提供逐标记指导。我们构建了一个三层对称语料库,涵盖文字推理转为语音、复杂声景中的事件推理,以及包含副语言线索的对话推理。该设计将跨模态蒸馏拓展至非语言事件、语调和对话上下文等范畴。在 MMSU、MMAU、BIG Bench Audio 和 MMAR 上的实验表明,X³-OPD 显著提升了音频基础推理能力和思维链质量,同时在领域偏移下较好保留了模型原有性能。
原文摘要 · Abstract (English)
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched textual inputs and verified answers. We further construct a three-tier symmetric corpus covering textual reasoning rendered into speech, audio-event reasoning grounded in complex acoustic scenes, and spoken-dialogue reasoning involving paralinguistic cues. This design extends cross-modal distillation beyond textually recoverable content to reasoning grounded in non-linguistic events, prosody, and conversational context. Experiments on MMSU, MMAU, BIG Bench Audio, and MMAR demonstrate that X$^3$-OPD substantially improves audio-grounded reasoning and chain-of-thought quality while largely preserving the model's existing capabilities under domain shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。