arXiv:2606.03116eess.AScs.AI2026-06被引 2

为语音指令生成设计动态评分标准,提升对齐评估的精准度与可解释性。

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

论文配图:AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
图 1 · 摘自论文原文
  • 基于动态评分项分解复杂指令,实现可验证的二元判断。
  • 在7920个样本上验证,零样本对齐检测显著优于现有方法。
  • 适合需要可解释奖励信号的音频生成强化学习研究者使用。

指令引导的音频生成技术快速发展,亟需可靠的对齐评估方法。现有自动化评估多依赖通用大语言模型的综合评分,难以解耦复杂指令、缺乏可解释性,且无法捕捉细微属性偏差。为此,我们提出一种动态评分基准评估范式,能自适应地将复杂音频描述分解为可独立验证的二元评分项。为严格评测该能力,我们构建了AnyAudio-Judge Bench,一个涵盖语音、声音、音乐和混合四类音频的双语基准,包含7,920个精心筛选的样本,含刻意设计的困难负例。同时,我们构建了一个包含10.5万条样本的大型数据集,附带显式的思维链(Chain-of-Thought)推理过程,用于训练专用评估器AnyAudio-Judge模型。通过结合监督微调(SFT)与组相对策略优化(GRPO)的训练流程,模型成功使其推理路径与评分机制对齐。大量实验表明,AnyAudio-Judge不仅显著提升了零样本对齐检测性能,相比当前最优基线,还提供了精确且可解释的奖励信号,大幅改善下游音频生成强化学习中的指令对齐效果。

原文摘要 · Abstract (English)

The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a novel dynamic rubric-based evaluation paradigm that adaptively decomposes complex audio captions into a variable number of independent, verifiable binary rubric items. To rigorously benchmark this capability, we propose the AnyAudio-Judge Bench, a comprehensive, bilingual benchmark comprising 7,920 meticulously curated samples across four diverse audio domains (speech, sound, music, and mixed), featuring deliberately constructed hard negatives. Furthermore, we construct a large-scale corpus of 105K samples with explicit Chain-of-Thought (CoT) rationales to train our dedicated evaluator, the AnyAudio-Judge model. By employing a training pipeline that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), our model successfully aligns its reasoning paths with the rubric-based scoring mechanism. Extensive experiments demonstrate that AnyAudio-Judge not only significantly enhances zero-shot alignment detection compared to state-of-the-art baselines, but also provides precise and interpretable reward signals that substantially improve instruction alignment in downstream reinforcement learning for audio generation.

音频生成指令对齐评估基准强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。