arXiv:2607.20454cs.CLcs.AI2026-07

首次系统评估前沿大模型响应漂移,发现所有模型都存在输出偏差。

Response drift across frontier large language models

论文配图:Response drift across frontier large language models
图 1 · 摘自论文原文
  • 10个主流大模型在62个跨领域问题上接受盲评,共2.9万次人工评估
  • 8个模型漂移率高达78%-81%,2个表现较好为47%-49%
  • 漂移模式因领域和问题而异,自动化指标无法捕捉人类判断

所有前沿大语言模型均存在响应漂移——输出与专家验证参考答案偏离,但其程度与结构尚未通过系统性人工评估厘清。本文报告了一项完全交叉的评估:47名来自不同地理区域的参与者,在盲态条件下对10个前沿大模型在62个跨领域问题上的输出进行评估,共获得29,140次独立判断。所有模型均出现漂移,但幅度差异显著:8个模型趋于统计上无差别的上限(78%-81%偏离),另有2个模型偏离更低(47%-49%)。漂移模式在六个领域和62个问题间存在差异,且上限模型间的成对相关系数超过r=0.85。自动化相似度指标解释的人类判断方差不足2%。结果表明,响应漂移在前沿大模型中普遍存在,其结构具有领域和问题依赖性,且仅可通过以人为中心的评估揭示。

原文摘要 · Abstract (English)

All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgements. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.

大模型评估响应漂移人类评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。