用评分标准提升大模型复杂指令理解能力,效果显著。
AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- 构建1600+条复杂指令的评分基准,支持精准评估
- 通过评分验证与奖励优化,指令遵循能力提升6.7%
- 适合想提升模型指令理解力的研究者和工程师
大型语言模型在多项任务上已取得显著进展,但对复杂、多轮及系统级指令的高级指令遵循(IF)仍面临挑战。缺乏高质量人工标注的基准和可靠可解释的奖励信号,制约了有效评估与训练。本文提出AdvancedIF,一个包含超过1,600个提示的综合性基准,配有专家设计的评分细则,用于评估模型在复杂、多轮及系统级指令下的表现。我们进一步提出RIFL(基于评分的指令遵循学习),一种新型后训练流程,结合评分生成、微调评分验证器与奖励塑造,实现高效的强化学习训练。大量实验表明,RIFL显著提升模型指令遵循能力,在AdvancedIF上实现6.7%的绝对性能增益,并在多个公开基准上表现优异。消融实验验证了各组件的有效性。本工作确立了评分体系在训练与评估高级指令遵循中的关键作用,为更强大可靠的AI系统铺平道路。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has led to impressive performance on a range of tasks, yet advanced instruction following (IF)-especially for complex, multi-turn, and system-prompted instructions-remains a significant challenge. Rigorous evaluation and effective training for such capabilities are hindered by the lack of high-quality, human-annotated benchmarks and reliable, interpretable reward signals. In this work, we introduce AdvancedIF (we will release this benchmark soon), a comprehensive benchmark featuring over 1,600 prompts and expert-curated rubrics that assess LLMs ability to follow complex, multi-turn, and system-level instructions. We further propose RIFL (Rubric-based Instruction-Following Learning), a novel post-training pipeline that leverages rubric generation, a finetuned rubric verifier, and reward shaping to enable effective reinforcement learning for instruction following. Extensive experiments demonstrate that RIFL substantially improves the instruction-following abilities of LLMs, achieving a 6.7% absolute gain on AdvancedIF and strong results on public benchmarks. Our ablation studies confirm the effectiveness of each component in RIFL. This work establishes rubrics as a powerful tool for both training and evaluating advanced IF in LLMs, paving the way for more capable and reliable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。