arXiv:2508.11818cs.SDcs.LG2025-08被引 8

让音频模型学会像人一样逐步推理,提升听觉理解能力。

Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding

  • 将音频问答数据转为带思考过程的链式训练样本
  • 在124万条数据上微调后,推理能力显著提升
  • 适合研究音频理解与多模态推理的学者

链式思维推理在大语言模型和视觉语言模型中表现优异,但在音频语言模型中的潜力尚未被充分探索。本文提出首个针对声音理解的链式思维评估基准AF-Reasoning-Eval,用于评测常识推理与细微选项区分能力。为构建训练语料,我们设计自动管道,将现有音频问答与分类数据转化为显式推理链,生成包含124万样本的AF-CoT-Train数据集。我们在Audio Flamingo系列模型上进行微调,结果显示在多个推理基准上性能显著提升,验证了链式思维微调在高级声音理解任务中的有效性。

原文摘要 · Abstract (English)

Chain-of-thought reasoning has demonstrated significant improvements in large language models and vision language models, yet its potential for audio language models remains largely unexplored. In this technical report, we take a preliminary step towards closing this gap. For better assessment of sound reasoning, we propose AF-Reasoning-Eval, a benchmark targeting common-sense reasoning and the ability to discriminate among closely related choices. To prepare training corpus for sound reasoning abilities, we propose automatic pipelines that transform existing audio question answering and classification data into explicit reasoning chains, yielding AF-CoT-Train with 1.24M samples. We study the effect of finetuning Audio Flamingo series on AF-CoT-Train and observe considerable improvements on several reasoning benchmarks, validating the effectiveness of chain-of-thought finetuning on advanced sound understanding.

音频理解链式思维多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。