提示词中的数字锚定会虚增大模型错误检测的评分,但实际定位错误能力没提升。
Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

- 设计测试框架ErrorBench,评估提示词对错误计数的影响。
- 数字锚定使F1分数最高虚增0.79分,真实定位精度提升不足0.04。
- 建议评测时避免预设错误数量,应同时报告定位精度指标。
计数为基础的F1分数被广泛用作大模型错误检测质量的代理指标,但本文揭示其可能因提示词设计导致虚假提升,称为F1膨胀。本文提出ErrorBench,一种受控的压力测试协议,用于检测提示词引发的计数偏差。在143个CoNLL-2014 M2风格篇章上,对六种主流大模型在五种提示条件下共4,290次响应进行评估。结果显示,在标准评分下,带锚定提示可导致最高0.79的F1膨胀,严格匹配下达0.96。使用官方ERRANT 3.0.0管道和多参考评分的100篇复制实验验证:六模型平均,从无锚定到带锚定提示,计数F1提升+0.21,而多参考ERRANT F0.5仅提升+0.04。高指令遵从性模型(如GPT/Claude)在测试中产生更多错误计数,而Gemini家族则较少。研究建议大模型校对与文档评审评估应避免预设错误数量,并报告与计数并行的跨度感知指标。
原文摘要 · Abstract (English)
Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。