arXiv:2601.16547cs.SDcs.AI2026-01ACL被引 6

用少量数据提升音频推理能力,让模型听懂声音背后的逻辑。

CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation

  • 用文本作内部教师,实时对齐音频与文本的推理过程。
  • 仅用8万条合成数据,显著缩小音频与文本任务表现差距。
  • 适合研究多模态对齐、高效训练的学者与工程师。

大型音频语言模型(LALMs)受到广泛关注。尽管基于文本大语言模型(LLMs),LALMs 常表现出知识与推理能力下降。我们假设该问题源于现有训练范式无法有效弥合特征空间中的声学-语义鸿沟。为此,我们提出 CORD,一种统一的对齐框架,实现在线跨模态自蒸馏。具体地,它在统一模型内对齐音频条件推理与文本条件推理。利用文本模态作为内部教师,CORD 在音频生成过程中进行多粒度对齐。在标记级别,采用带重要性加权的策略反KL散度,优先关注早期且语义关键的标记;在序列级别,引入基于判别器的全局奖励,通过分组相对策略优化(GRPO)优化完整推理轨迹。多个基准测试的实证结果表明,CORD 持续提升音频条件推理能力,并仅用8万条合成训练样本即显著缩小音频-文本性能差距,验证了其有效性与数据高效性。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have garnered significant research interest. Despite being built upon text-based large language models (LLMs), LALMs frequently exhibit a degradation in knowledge and reasoning capabilities. We hypothesize that this limitation stems from the failure of current training paradigms to effectively bridge the acoustic-semantic gap within the feature representation space. To address this challenge, we propose CORD, a unified alignment framework that performs online cross-modal self-distillation. Specifically, it aligns audio-conditioned reasoning with its text-conditioned counterpart within a unified model. Leveraging the text modality as an internal teacher, CORD performs multi-granularity alignment throughout the audio rollout process. At the token level, it employs on-policy reverse KL divergence with importance-aware weighting to prioritize early and semantically critical tokens. At the sequence level, CORD introduces a judge-based global reward to optimize complete reasoning trajectories via Group Relative Policy Optimization (GRPO). Empirical results across multiple benchmarks demonstrate that CORD consistently enhances audio-conditioned reasoning and substantially bridges the audio-text performance gap with only 80k synthetic training samples, validating the efficacy and data efficiency of our on-policy, multi-level cross-modal alignment approach.

多模态对齐音频推理自蒸馏小样本训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。