用视觉模型推理能力教会音频模型逐步思考,提升听觉理解效果。
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- 从视觉模型迁移分步推理能力到音频模型
- 在音视频问答任务中显著提升未知场景表现
- 适合需要强听觉推理的多模态研究者
尽管大型音频-语言模型(LALM)已实现顶尖的音频理解能力,其在复杂声景中的推理能力仍不及大型视觉-语言模型(LVLM)。与视觉领域相比,主要瓶颈在于缺乏大规模链式思维(CoT)音频数据来训练模型进行逐步推理。为克服这一数据和模态差距,我们提出SightSound-R1,一种跨模态推理蒸馏框架,将更强的LVLM教师模型的高级推理能力,迁移到较弱的LALM学生模型上,使用相同的音视频问答(AVQA)数据集。SightSound-R1包含三个核心步骤:(i) 测试时扩展,生成以音频为中心的链式思维(CoT),(ii) 基于音频的验证,过滤幻觉内容,(iii) 采用监督微调(SFT)结合组相对策略优化(GRPO)的蒸馏流程。结果表明,SightSound-R1不仅在域内AVQA测试集上提升性能,在未见听觉场景和问题上也表现更优,优于预训练模型及仅依赖标签蒸馏的基线模型。因此,我们得出结论:视觉推理可有效迁移到音频模型,并通过大量音视频数据实现扩展。
原文摘要 · Abstract (English)
While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio data to teach LALM stepwise reasoning. To circumvent this data and modality gap, we present SightSound-R1, a cross-modal distillation framework that transfers advanced reasoning from a stronger LVLM teacher to a weaker LALM student on the same audio-visual question answering (AVQA) dataset. SightSound-R1 consists of three core steps: (i) test-time scaling to generate audio-focused chains of thought (CoT) from an LVLM teacher, (ii) audio-grounded validation to filter hallucinations, and (iii) a distillation pipeline with supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for the LALM student. Results show that SightSound-R1 improves LALM reasoning performance both in the in-domain AVQA test set as well as in unseen auditory scenes and questions, outperforming both pretrained and label-only distilled baselines. Thus, we conclude that vision reasoning can be effectively transferred to audio models and scaled with abundant audio-visual data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。