arXiv:2607.00304cs.LGcs.AI2026-07综述

实证揭示大模型评估中偏见与可靠性的权衡关系

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

  • 扩展至11种评估条件,系统测量耦合度、策略多样性与小样本可靠性
  • 低耦合条件噪声高(CV(N=5)>1.0),高耦合条件噪声低(CV(N=5)<0.16)
  • GPT-4o出现异常模式,或因API版本漂移,适合评估系统研究者参考

该研究拓展了对大模型评估中偏见-可靠性权衡的实证基础,从此前单一研究的5个条件扩展至11个。在所有11个条件中测量了评估器耦合度(gamma)和策略多样性(H),其中9个具备有效权重向量;对7个具备足够种子数的条件(N≥5)测量了小样本可靠性(CV(N=5))。五个条件获得完整的(gamma, H, CV)三元组。数据证实该权衡:当评估器耦合度低(gamma < 0.2)时,测量噪声高(CV(N=5) > 1.0);而强耦合(gamma > 0.9)条件下,噪声极低(CV(N=5) < 0.16)。相关系数r(H, gamma) = -0.989(n=5,排除GPT-4o条件)表明耦合度抑制策略多样性。四个GPT-4o条件显示gamma=0.000且H=1.000,我们归因于2026年6月GPT-4o API的版本漂移。无任何条件同时满足{gamma < 0.2, CV(N=5) < 0.3}。所有条件的度量结果已作为标准化基准数据集发布,供评估器对比。

原文摘要 · Abstract (English)

The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma < 0.2) exhibit high measurement noise (CV(N=5) > 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.

大模型评估偏见-可靠性评估系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。