arXiv:2604.16617cs.CVcs.MM2026-04被引 1

用单模态模型生成音视频推理数据,提升多模态推理能力

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

论文配图:AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
图 1 · 摘自论文原文
  • 用专精视觉和音频的模型分别生成推理轨迹,再由大模型合并
  • 3B和7B模型在多个音视频基准上达到领先水平
  • 适合想训练多模态推理模型的研究者和工程师

近期推理模型在文本领域取得显著进展,但将其能力迁移到音视频等多模态场景仍面临挑战,主要受限于目标多模态组合中高质量推理数据的匮乏。为此,我们提出AVRT框架,通过单模态教师模型生成高质量音视频推理轨迹。首先利用专精于视觉或音频的模型分别生成独立的视觉与音频推理轨迹,再通过一个大语言模型合并器进行融合。所生成的多模态推理轨迹用于监督微调(SFT)冷启动阶段,先将目标模型适配至音视频推理任务,随后在更大规模数据上进行强化学习第二阶段训练。在七个音视频及音频基准上评估,我们的3B和7B参数模型在同类规模模型中表现最佳,超越OmniBench、DailyOmni(音视频)及MMAR(纯音频)等基线,表明跨模态训练也能泛化到单模态任务,并建立了一条新的多模态推理模型训练范式。

原文摘要 · Abstract (English)

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, in part because of the limited availability of high-quality reasoning data in targeted multimodal combinations. To address this problem, we introduce AVRT, a novel framework that generates high-quality audio-visual reasoning traces from single-modality teacher models. We generate independent vision- and audio-reasoning traces via models specialized to reason over their respective modalities and merge the resulting traces with an LLM merger model. The resulting multimodal traces are used in a supervised fine-tuning (SFT) cold start to adapt the target model to audio-visual reasoning traces first, before training it in a second reinforcement learning stage on larger-scale data. Evaluated on seven audio-visual and audio benchmarks, our 3B and 7B parameter models achieve state-of-the-art results among models of comparable size including OmniBench and DailyOmni for audio-visual and MMAR for audio-only reasoning, showing that cross-modal training also transfers to single-modality tasks and establishing a new training pipeline for multimodal reasoning models.

多模态推理音视频理解知识蒸馏大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。