arXiv:2506.12935cs.CLcs.MM2025-06EMNLP被引 20

用强化学习提升音频语言模型的逻辑推理能力

SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models

  • 设计规则化强化学习算法,驱动模型理解音频与文本逻辑关系
  • 在6446个标注样本上训练,性能超越现有最佳方法
  • 适合研究音频智能、多模态推理的学者与开发者

尽管大语言模型展现出强大推理能力,但其向音频模态的拓展,尤其在大型音频-语言模型(LALMs)中的应用仍不充分。本文提出一套系统性解决方案:构建包含6,446个音频-文本标注样本的SoundMind数据集,聚焦复杂推理任务;在此基础上,提出SoundMind-RL——一种基于规则的强化学习算法,用于赋予音频-语言模型稳健的跨模态推理能力。通过在Qwen2.5-Omni-7B上使用SoundMind数据集进行微调,模型在SoundMind基准测试中显著优于当前最先进方法,验证了高质量推理数据与专用强化学习技术结合的有效性。该工作推动了语言模型在听觉智能方面的进展。代码与数据集已公开于https://github.com/xid32/SoundMind。

原文摘要 · Abstract (English)

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic approach that involves a capable base model, high-quality reasoning-oriented audio data, and effective training algorithms. In this work, we present a comprehensive solution for audio logical reasoning (ALR) tasks: we introduce SoundMind, a dataset of 6,446 audio-text annotated samples specifically curated to support complex reasoning. Building on this resource, we propose SoundMind-RL, a rule-based reinforcement learning (RL) algorithm designed to equip audio-language models with robust audio-text reasoning capabilities. By fine-tuning Qwen2.5-Omni-7B on the proposed SoundMind dataset using SoundMind-RL, we achieve strong and consistent improvements over state-of-the-art baselines on the SoundMind benchmark. This work highlights the benefit of combining high-quality, reasoning-focused datasets with specialized RL techniques, and contributes to advancing auditory intelligence in language models. The code and dataset introduced in this work are publicly available at https://github.com/xid32/SoundMind.

音频推理强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。