arXiv:2608.02966cs.CL2026-08

让大模型的错误选项也能被量化分析,揭示其真实能力水平。

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

  • 构建可解析错误选项的统计模型,同时估计模型能力与题目特性。
  • 错误答案信息量比正确答案多101%,单独用错题就能准确评估模型能力。
  • 只需41个精选题目即可还原全库排名,效率提升770倍,适合高效评测。

现有大语言模型多选题评测仅关注答对与否,将所有错误回答等同处理,忽视了错误选项中蕴含的系统性行为信息。本文提出LLM名义反应模型(LLM-NRM),一种选项感知的心理测量框架,能建模完整作答分布,联合估计模型能力与题目层级特征,分离出模型响应锐度、位置偏好及难度依赖的退化行为。在189个大模型与31,554道题目(来自14个基准)上,该模型预测未见模型-题目交互更精准,能力估计与外部人类偏好排行榜(Arena.ai Elo)的斯皮尔曼相关系数达0.920,显著优于传统方法。错误选项贡献的费雪信息量比正确答案高出101%;仅用错误回答即可实现0.943的相关性,恢复完整信息。学习到的题目参数支持高效评测:仅需41个样本题目即可保留全库排名,肯德尔相关为0.85,效率提升770倍。结果表明,错误答案承载着独特且有价值的测量信息,而非等价失误。

原文摘要 · Abstract (English)

Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.

心理测量大模型评测错误分析高效测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。