arXiv:2606.09118cs.AI2026-06中稿 · ACL被引 1

用专家设计的评分标准提升大模型指令遵循和代理行为。

ComplexConstraints and Beyond: Expert Rubrics for RLVR

  • 用专家制定的评分规则作为统一评估与强化学习奖励信号。
  • 在ComplexConstraints上训练4B模型,关键指标提升15.5个百分点。
  • 效果可迁移至未见过的外部基准,适合训练复杂任务模型。

评估方法常落后于大模型能力。程序化验证基准仅覆盖表层约束,而真实世界中的指令遵循与代理工作流需评估语义、上下文及策略依赖行为。本文研究专家定制评分标准作为跨复杂指令遵循与企业级代理任务的统一评估与强化学习奖励机制。识别出影响奖励质量的关键设计:最大可行原子性、意图感知准则设计与大模型裁判校准。提出ComplexConstraints,一个专家制定的指令遵循套件,包含75个公开提示的基准集(1,559条评分标准)和1,000个独立提示的训练集,每提示含10-40条原子标准。实证表明,评分奖励在固定任务数据集(如ComplexConstraints)和状态化强化学习环境(如CoreCraft)中均有效提升训练。在ComplexConstraints上训练4B模型,使保留测试集的平均标准通过率提升15.5个百分点,接近60倍大的Qwen3模型未训练基线水平(仅差0.5个百分点),且性能迁移至未见基准:AdvancedIF提升8.4个百分点,MultiChallenge提升10.1个百分点。在CoreCraft中,评分奖励强化学习同样在分布外基准上实现转移增益:BFCL+4.5 pp,tau^2-Bench+7.4 pp,Toolathlon+6.8 pp。结果表明,专家撰写的评分标准能提供有效的评估目标与可扩展的奖励信号,显著提升大模型的指令遵循与代理行为。

原文摘要 · Abstract (English)

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic workflows require judging semantic, contextual, and policy-dependent behavior. We study expert-curated rubric-based evaluation as a unified mechanism for measurement and reinforcement-learning rewards across two settings: complex instruction following and enterprise agentic tasks. We identify rubric-design choices that affect reward quality, including maximum viable atomicity, intent-aware criterion design, and LLM-judge calibration. We introduce ComplexConstraints, an expert-curated instruction-following suite comprising a public 75-prompt benchmark with 1,559 rubric criteria and a disjoint 1,000-prompt training set, with 10-40 atomic criteria per prompt. Empirically, rubric rewards improve training in both fixed task datasets, such as ComplexConstraints, and stateful RL environments, such as CoreCraft. Training a 4B model on ComplexConstraints improves mean criterion pass rate by +15.5 pp on a held-out split, bringing it within 0.5 pp of the untrained baseline of a roughly 60x larger Qwen3 model, and the gains transfer to external benchmarks the model never saw during training: +8.4 pp on AdvancedIF and +10.1 pp on MultiChallenge. In CoreCraft, rubric-reward RL likewise transfers to out-of-distribution benchmarks (+4.5 pp BFCL, +7.4 pp tau^2-Bench, +6.8 pp Toolathlon). These results show that expert-authored rubrics provide effective evaluation targets and scalable reward signals for improving LLM instruction following and agentic behavior.

大模型评估强化学习指令遵循专家评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。