arXiv:2606.05122cs.CL2026-06被引 2

无需训练,大模型可自评回答质量。

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

论文配图:Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
图 1 · 摘自论文原文
  • 通过校准强化学习与掩码蒸馏,唤醒模型潜藏的自评能力。
  • 仅用160个样本,自评校准效果优于基准方法31倍。
  • 自评结果稳定且可迁移,不依赖特定评分者偏好。

大型语言模型越来越多地由其他模型进行评估,这引发了一个自然问题:模型能否预测其自身输出将被评分者的打分?我们发现这种能力在未经专门训练前就已存在:通过少量示例提示,基础模型即可在三个基准测试上,对开放式回答的多维度质量评分做出显著高于随机水平的预测。我们提出自评估激发(SEE)方法,通过一个包含校准强化学习和掩码蒸馏的短周期,提升答案质量并精准预测评分者打分,同时保持答案不变。仅使用160个独特样本(约为强化学习基线的1/31),SEE在三个基准测试上提升了保留数据的校准度,且保持答案质量。激发后的自评能力集中于模型自身词元分布内,并在未参与训练的评分者间保持稳定,表明这是一种可迁移的质量认知,而非单一评分者的偏好。这些结果将评分对齐的自评估重构为一个激发而非获取的问题。

原文摘要 · Abstract (English)

Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an external judge's multi-attribute quality scores on open-ended responses well above chance across three benchmarks. We introduce Self-Evaluation Elicitation (SEE), a method that surfaces this latent ability through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while leaving the answer untouched. From 160 unique examples, roughly 31x fewer than a reinforcement learning baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality. The elicited self-evaluation is sharply localized within the model's own token distribution and stable across judges it was never trained against, indicating a transferable notion of quality rather than a single judge's preference. These results reframe judge-aligned self-evaluation as a problem of elicitation rather than acquisition.

自评估模型校准少样本可迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。