arXiv:2608.02831cs.SDcs.CL2026-08被引 1

用可自演化的评分标准,让AI听懂音频并给出合理推理。

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

论文配图:Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
图 1 · 摘自论文原文
  • 基于音频波形动态生成评分标准,随模型表现自动调整。
  • 在三个基准上显著超越现有方法,推理长度趋于稳定。
  • 适合需要精准音频理解与逻辑推理的任务研究者。

音频推理对机器理解声学世界至关重要。基于可验证奖励的强化学习能激发此类推理,但现有奖励设计存在局限:结果导向奖励仅监督最终答案,导致模型忽略音频细节;过程导向奖励依赖粗糙、手工固定的标准,无法随问题变化而调整,且随着策略提升逐渐失效。不同问题对感知或多步推理的需求各异,静态标准难以适应。因此,需细粒度、音频相关且可自适应的奖励机制。为此,我们提出AudioRubrics框架,通过原始波形合成每样本的评分标准,并根据模型自身推演结果,动态重构与重加权评价准则,提供持续学习信号,针对当前策略弱点进行优化。在三个音频推理基准上的综合评估显示,AudioRubrics显著优于多种开源及训练基线。分析表明,性能提升与评分生成器和评判器能力正相关,且模型收敛至稳定推理长度,避免退化或无限增长。音频感知能力的提升也验证了以声学证据锚定监督的有效性。项目页面见https://audiorubrics.github.io。

原文摘要 · Abstract (English)

Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and let the model reach it without attending to the audio, whereas process-based rewards score the reasoning itself but rely on coarse, hand-crafted, and fixed criteria that neither adapt to each question nor stay grounded in the acoustic evidence. Moreover, questions differ in what they demand, with some hinging on perception and others on multi-step reasoning, and any static criterion weakens as the policy improves. Supervising the reasoning process with fine-grained, audio-grounded, and adaptive rewards is therefore crucial, yet challenging since such rewards are impractical to design by hand for every sample. To this end, we introduce AudioRubrics, a reinforcement learning framework that supervises audio reasoning with self-evolving, audio-grounded rubric rewards. AudioRubrics synthesizes per-sample rubrics from the raw waveform and, conditioned on the model's own rollouts, regenerates and reweights criteria per group, supplying a continuous learning signal that keeps targeting the current policy's weaknesses as static criteria saturate. Comprehensive evaluations across three audio reasoning benchmarks reveal that AudioRubrics substantially outperforms a wide range of open-source and training-based baselines. Furthermore, our analysis shows that the gains scale with the capability of the rubric generator and judge, and AudioRubrics converges to a stable reasoning length that avoids both degenerate collapse and unbounded growth. The improvement in audio perception further demonstrates the effectiveness of anchoring supervision in the acoustic evidence. Our project page is available at https://audiorubrics.github.io.

音频推理强化学习自演化评分标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。