arXiv:2606.20093cs.CL2026-06被引 1

测试发现大模型在修正自己文本时无自偏见,反而更挑剔错误。

Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship

  • 用可验证的指令遵循任务测试模型是否偏好自己生成内容
  • 作者身份下拒绝有效修改率与中立模型基本一致(差值-5.1个百分点)
  • 97%拒绝理由是找错而非偏好,适合研究模型自我修正机制的人看

大型语言模型越来越多地参与文本审查与修改,包括对自己的输出进行修订。已有研究表明,当作为评判者时,模型存在偏好自身生成内容的自偏见,这引发疑问:模型是否会抗拒对自己写作的有效修正?本研究在一种真实作者身份设定下测试此问题——使用IFEval的确定性验证器判断修正是否有效。模型先生成草稿,由官方IFEval检查器确认其违反约束,并验证候选修改有效;随后模型以真实作者身份或中立身份决定接受或拒绝该修改。在四个中等规模模型家族、85组作者与中立模型对比中,未发现可检测的自偏好:作者拒绝经验证有效的修改率与中立模型几乎相同(差距-5.1个百分点,95%置信区间[-12.9, +2.7])。小规模初步结果中的自怀疑倾向未在大规模实验中复现。唯一稳健的定性发现是:当作者拒绝有效修改时,97%的理由为‘找错’,而非偏好,说明拒绝行为源于质量判断而非自恋。样本量下小于约13个百分点的效应无法排除。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly review and revise text, including their own. A documented self-preference bias (models favoring their own generations when acting as judges) raises the question of whether models also resist valid corrections to their own writing. We test this in a setting where "valid" is decided not by another model but by a deterministic verifier: instruction-following revision on IFEval. A model writes a draft; the official IFEval checker confirms the draft violates a constraint and that a candidate edit fixes it; the model then accepts or rejects that edit either as the genuine in-context author or as a fresh model that sees the draft neutrally. Across four mid-tier model families and 85 author-versus-fresh comparisons, we find no detectable self-preference: authors reject verified-good fixes to their own drafts at essentially the same rate as fresh models judging the same drafts (gap -5.1 pp, 95% CI [-12.9, +2.7]). A self-skepticism hint from a smaller pilot did not replicate at scale. The one robust observation is qualitative: when authors do reject a verified-good fix, 97% of their stated reasons are flaw-catching rather than preference, that is, about the character of rejections, not an elevated rate. Effects smaller than ~13 pp cannot be excluded at this sample size.

自偏见模型修正可验证性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。