arXiv:2507.05598cs.CLcs.AI2025-07综述被引 1

用自评与精修框架提升大模型指令遵循能力,低成本高效保质。

Self-Review Framework for Enhancing Instruction Following Capability of LLM

  • 通过解析指令结构实现分步自评与选择性修正
  • 小数据量下达成媲美GPT-4o-mini的指令遵循效果
  • 适合资源有限但需高精度指令执行的场景

为提升大语言模型对格式与指令约束的遵循能力,现有方法多依赖高性能模型生成高质量数据,但其常无法一次性满足复杂指令。迭代修订虽有效,却带来高昂成本。现有基于评估工具的方法又易因过度修订导致输出质量下降。为此,我们提出Re5框架:从用户指令中提取任务与约束组件,进行结构化评估以防止错误累积,并实施细粒度约束专项评估后选择性修订,确保精准改进且不损质量。最终高质量输出用于对齐调优,形成以数据为中心的迭代优化闭环。实验表明,仅用少量数据,Re5在指令遵循性能上接近使用GPT-4o-mini生成数据训练的模型,同时保持响应质量,相比未修订初始结果胜出64.24%。验证了Re5在极低外部监督下高效提升指令遵循能力的有效性。

原文摘要 · Abstract (English)

Various techniques have been proposed to improve large language models (LLMs) adherence to formatting and instruction constraints. One of the most effective approaches involves utilizing high-quality data generated by powerful models. However, such models often fail to fully comply with complex instructions in a single generation. To address this limitation, iterative revision methods have been introduced. Nevertheless, as the number of data points and revision iterations increases, the associated monetary costs grow significantly. As a resource-efficient alternative, methods have been proposed that leverage high-performance evaluation tools to compensate for the limited self-evaluation capabilities of open-source LLMs. However, these approaches often lead to a degradation in output quality due to excessive revision. To overcome these challenges, we propose Re5, a self-evaluation and revision framework designed to enhance instruction-following performance while preserving the quality of the generated content. Re5 extracts task and constraint components from user instructions, performs structural evaluations to prevent error accumulation, and applies fine-grained constraint-specific content evaluations followed by selective revisions. This process ensures precise and quality-preserving improvements. The final high-quality outputs are used for alignment tuning, enabling long-term alignment improvements through a data-centric iterative refinement loop. Experimental results demonstrate that Re5 achieves instruction-following performance comparable to models trained on data generated by GPT-4o-mini, a high-performance model, even with a small amount of data while maintaining response quality with a 64.24%-win rate over the non-revised initial responses. These results validate Re5 as an efficient and effective solution for enhancing instruction adherence with minimal external supervision.

指令遵循自评修正大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。