arXiv:2602.02589cs.AIcs.LG2026-02综述被引 3

用自动生成的同行评审,实现无需人工的智能评估。

PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review

  • 模型自主生成问题并联网回答,互评打分,全程无监督。
  • 在12个商用模型上稳定区分性能,与Elo评分高度一致。
  • 可发现身份和呈现偏差,适合开放世界模型评测。

大语言模型评估通常依赖人工构建的基准、参考答案或单一模型判断,这些方法难以扩展、易过时,且与依赖网络检索和融合的开放世界应用不匹配。我们提出PeerRank,一种完全自动化的端到端评估框架:模型自主生成评估任务,使用实时网络信息进行类别限定的回答,评判其他模型的回复,并将密集的同行评估聚合为相对性能估计,无需人工干预或标准答案。PeerRank将评估视为多智能体过程,每个模型对称地承担任务设计者、回答者和评估者角色,同时消除偏见。在覆盖12个商用模型和420个自动生成问题的大规模研究中,PeerRank产生稳定且具有区分力的排名,并揭示了可测量的身份和呈现偏差。排名具有鲁棒性,平均同伴评分与Elo评分高度一致。我们在TruthfulQA和GSM8K上进一步验证,同伴评分与客观准确率相关。结果表明,结合选择性网络接地的有偏意识同行评估,可突破静态人工基准的局限,实现开放世界大模型评估的规模化。

原文摘要 · Abstract (English)

Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deployments that depend on web retrieval and synthesis. We introduce PeerRank, a fully autonomous end-to-end evaluation framework in which models generate evaluation tasks, answer them with category-scoped live web grounding, judge peer responses and aggregate dense peer assessments into relative performance estimates, without human supervision or gold references. PeerRank treats evaluation as a multi-agent process where each model participates symmetrically as task designer, respondent, and evaluator, while removing biased judgments. In a large-scale study over 12 commercially available models and 420 autonomously generated questions, PeerRank produces stable, discriminative rankings and reveals measurable identity and presentation biases. Rankings are robust, and mean peer scores agree with Elo. We further validate PeerRank on TruthfulQA and GSM8K, where peer scores correlate with objective accuracy. Together, these results suggest that bias-aware peer evaluation with selective web-grounded answering can scale open-world LLM assessment beyond static and human curated benchmarks.

模型评估自主评测网络接地同行评审

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。