arXiv:2509.12275cs.SDcs.AI2025-09被引 2

通过错因引导的课程学习,提升音频问答模型推理能力

Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering

  • 按难易度组织数据,错题优先训练
  • 在MMAR上达到64.30%新高,MMAU-mini达73.80%
  • 适合想提升音频推理性能的研究者

随着大音频语言模型(LALMs)的快速发展,音频问答(AQA)成为需要精细音频理解与复杂推理的挑战性任务。当前方法主要依赖通过字幕或推理轨迹构建新数据集,但高质量现有AQA数据仍被低估。为此,我们提出Omni-CLST:一种错因感知的课程学习框架,结合引导式选择性思维链。该框架通过两项关键策略高效利用已有高质量数据:基于错误感知的课程安排,按难度组织样本;以及引导式思维丢弃机制,聚焦于困难案例的推理。实验表明,Omni-CLST在MMAU-mini上达到73.80%,在新基准MMAR上创下64.30%的最新纪录,展现出在多模态音频语言理解中的强大泛化能力。

原文摘要 · Abstract (English)

With the rapid progress of large audio-language models (LALMs), audio question answering (AQA) has emerged as a challenging task requiring both fine-grained audio understanding and complex reasoning. While current methods mainly rely on constructing new datasets via captioning or reasoning traces, existing high-quality AQA data remains underutilized. To address this, we propose Omni-CLST, an error-aware Curriculum Learning framework with guided Selective Chain-of-Thought. The framework efficiently leverages existing high-quality dataset through two key strategies: an error-aware curriculum that organizes samples by difficulty, and a guided thought dropout mechanism that focuses reasoning on challenging cases. Experiments show that Omni-CLST achieves 73.80% on MMAU-mini and a new state of the art of 64.30% on MMAR, demonstrating robust generalization in multimodal audio-language understanding.

音频问答课程学习推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。