arXiv:2509.19880cs.CLcs.AI2025-09EMNLP被引 2

用模型自答作为参考,提升大模型评估的可靠性

Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation

  • 让模型用自己生成的答案做评判基准
  • 自参考策略使生成与判断能力相关性提升显著
  • 适合需要可靠模型筛选的评估场景

LLM作为裁判的评估框架日益流行,但关于模型生成与判断能力之间关系的研究结果仍不一致。我们通过跨11个模型和21个多样化任务的系统性数据集与实例级分析,发现尽管两种能力依赖相同底层知识,但相关性较弱,主要源于大模型对被评判内容的敏感性。为此,我们提出一种自参考引导的评估策略,利用模型自身回答作为参考标准。该方法显著增强生成与判断能力间的相关性,为对齐这两项技能提供实用路径,并在评估任务中成为可靠的模型选择代理。

原文摘要 · Abstract (English)

LLM-as-Judge frameworks are increasingly popular for AI evaluation, yet research findings on the relationship between models' generation and judgment abilities remain inconsistent. We investigate this relationship through systematic dataset- and instance-level analyses across 11 models and 21 diverse tasks. Despite both capabilities relying on the same underlying knowledge, our analyses reveal they are only weakly correlated, primarily due to LLMs' sensitivity to the responses being judged. To address this, we propose a self-reference-guided evaluation strategy that leverages a model's own answers as references. This approach significantly strengthens the correlation between generation and judgment abilities, offering a practical path to align these skills and providing a reliable proxy for model selection in evaluation tasks.

大模型评估自参考模型筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。