不依赖音频也能提升语音大模型性能,关键靠文本推理优化
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
- 用强化学习在纯文本数据上微调语音大模型
- 在多个音频问答任务上达到最新最优成绩
- 发现文本推理能力提升可反哺音频理解
我们提出Omni-R1,通过强化学习方法GRPO,在音频问答数据集上微调近期多模态大模型Qwen2.5-Omni。该方法在MMAU和MMAR最新基准测试中取得新SOTA表现,于声音、音乐、语音及整体平均类别上,均在Test-mini与Test-full两个划分下达到最高准确率。为理解性能提升来源,我们对比了含音与不含音的模型,发现大部分性能提升源于更优的文本推理能力。更意外的是,仅在纯文本数据上微调,也能有效提升音频相关任务的表现。
原文摘要 · Abstract (English)
We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU and MMAR benchmarks. Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。