arXiv:2506.22157cs.CL2025-06ACL被引 5

让AI学会写能真正改进答案的批评,提升生成质量。

Training Language Model to Critique for Better Refinement

  • 用改进结果反向训练批评模型,让好批评带来好改写。
  • 在5个任务中显著优于传统方法,改写后效果提升明显。
  • 适合想优化大模型反馈能力的研究者和开发者。

大型语言模型在评估与批评方面表现出色,但关于何种批评最有助于改进结果、如何生成此类批评的研究仍有限。为此,我们提出一种名为精炼导向批评优化(RCO)的新框架,通过精炼信号训练批评模型。该框架采用反馈循环:批评模型生成的批评指导执行模型改进回应,批评有效性(CU)量化改进程度,作为批评模型的奖励信号。通过聚焦推动实质改进的批评,RCO无需直接进行批评偏好判断,确保有效批评获得奖励。我们在对话生成、摘要、问答、数学推理和代码生成五个任务上评估了RCO,结果表明其在批评质量和改写效果上均显著优于传统方法和开源模型。本研究贡献包括提出基于改写偏好监督的新范式及全面实验验证该方法的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. However, limited research has explored which types of critiques are most effective for improving model responses or how to generate such critiques. To address this gap, we introduce \textbf{R}efinement-oriented \textbf{C}ritique \textbf{O}ptimization (RCO), a novel framework designed to train critic models using refinement signals. RCO uses a feedback loop where critiques, generated by the critic model, guide the actor model in refining its responses. The critique utility (CU) quantifies the effectiveness of these refinements, serving as the reward signal for training the critic model. By focusing on critiques that lead to better refinements, RCO eliminates the need for direct critique preference assessment, ensuring that critiques driving meaningful improvements are rewarded. We evaluate RCO across five tasks, i.e., dialog generation, summarization, question answering, mathematical reasoning, and code generation, and show that it significantly outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes. Our contributions include the introduction of RCO, a novel supervision scheme based on refined response preferences, and comprehensive experimental results that highlight the method's effectiveness in enhancing LLM critique-refinement loops.

大模型批评优化生成改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。