arXiv:2607.02668cs.CL2026-07

让大模型生成与自检更一致,提升输出可靠性。

Improving LLMs via Validator-to-Generator Alignment

论文配图:Improving LLMs via Validator-to-Generator Alignment
图 1 · 摘自论文原文
  • 用频率修正的评分机制解决生成与验证不一致问题
  • 在IFEval和HumanEval上相关性提升最高达27个百分点
  • 适合关注模型输出一致性与可信度的研究者

大语言模型存在不一致性:不同提示或引入无关信息可能导致输出突变。生成-验证器(G-V)差距是这一现象的表现之一,即模型生成的回答在要求自检时被判定为无效。本文提出一种基于频率修正的新型G-V一致性形式,解决了生成器因先验低频而低估有效文本的问题。在理性多答案回答模型下,频率修正后的生成评分能自然实现验证器一致性。提出的 extit{ cpaname}( cpa)方法通过训练目标实现该一致性,在真实大模型上显著提升G-V一致性与生成性能,在IFEval和HumanEval上相关性最高提升27个百分点,同时保持所有任务中验证器质量不变。

原文摘要 · Abstract (English)

Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, \emph{\FCPAname} (\FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with \FCPA{} substantially improves both G-V consistency and generator performance over prior methods, with gains of up to $+27$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.

大模型一致性生成质量验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。