发现大模型错误严重性分布差异,揭示误差本质不只看数量。
ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

- 构建10,000条查询的连续严重性评分基准,按8个领域5级难度评估
- 21个开源模型中85对在准确率相同时严重性指数差异显著(95%置信区间不重叠)
- 高严重性错误多为虚构内容,低严重性多为检索错误,且随模型大小变化
在相同准确率下,开源大语言模型的错误严重性分布存在显著差异,这一差异无法通过单一误差率捕捉。现有幻觉评测仅统计错误数量并视为等价,但错误日期与虚构判例的严重性相差数个量级。我们提出Errorquake-10k,一个包含10,000个查询的基准,对每个响应在0–4的连续严重性尺度上评分,覆盖8个领域和5个难度层级,并为21个开源模型拟合其严重性分布。对每模型估计严重性分布指数b(Gutenberg-Richter上尾斜率),并给出95%自举置信区间。结果显示,在210组模型对中,85组在匹配准确率(|Δε| < 0.05)下,其b的95%置信区间无重叠(如deepseek-v3.2 vs. ministral-14b,ε=0.586,Δb=0.47)。三评者人工验证研究(共519项)确认测量可靠性(ICC(2,k=3)=0.85)、LLM评判排名有效性(rho=0.89),以及人类数据上的密集模型规模相关性(rho_s=-0.86)。我们证明非可约性定理:严重性分布与误差率信息上不冗余(I(b; model | ε) = 1.56 bit;64.5%的跨模型b方差无法由ε解释)。严重性机制分类学(kappa=0.83)显示,错误类型随严重性分层:低严重性错误主要为检索错误(71%);高严重性错误主要为虚构内容(39%),且该构成随模型规模变化(p<0.0001)。严重性分布应与准确率一同报告,其承载误差率无法传达的区分性信息。
原文摘要 · Abstract (English)
At matched accuracy, open-weight LLMs differ substantially in the shape of their error severity distribution -- a difference invisible to the scalar error rate. Hallucination benchmarks report a single error count and treat all errors as equivalent, yet a wrong date and a fabricated court ruling differ by orders of magnitude. We introduce Errorquake-10k, a 10,000-query benchmark scoring each response on a continuous 0-4 severity scale across 8 domains and 5 difficulty tiers, and we fit per-model severity distributions for 21 open-weight models. For each model we estimate a severity distribution index (b, the Gutenberg-Richter upper-tail slope) with 95% bootstrap confidence intervals. Headline: across the 210 model pairs, 85 have disjoint 95% b confidence intervals at matched accuracy (|Delta epsilon| < 0.05) on human-consensus scoring, e.g. deepseek-v3.2 vs. ministral-14b at epsilon = 0.586 and Delta b = 0.47. A 519-item three-rater human validation study confirms measurement reliability (ICC(2,k=3) = 0.85), validates the LLM-judge ranking (rho = 0.89), and confirms the dense-model scaling correlation on human data (rho_s = -0.86). We prove a Non-Reducibility Theorem showing that severity profile and error rate are informationally non-redundant (I(b; model | epsilon) = 1.56 bits; 64.5% of cross-model b variance is unexplained by epsilon). A severity mechanism taxonomy (kappa = 0.83) reveals that error type shifts categorically with severity: low-severity errors are retrievals (71%); high-severity errors are fabrications (39%) -- and this composition differs by model size (p < 0.0001). Severity distribution should be reported alongside accuracy; it carries discriminative information that the error rate cannot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。