arXiv:2608.30372cs.CL2026-08中稿 · EMNLP

用概率分布分析大模型答题表现,自动识别题目质量缺陷。

Auditing MCQA Benchmarks through Probability Landscapes

论文配图:Auditing MCQA Benchmarks through Probability Landscapes
图 1 · 摘自论文原文
  • 通过模型预测概率分布构建评测基准的全局画像。
  • 在四个基准上发现模型置信度与干扰项竞争差异。
  • 噪声注入法可定位需人工复核的可疑题目。

随着大语言模型快速进步,标准多选题问答(MCQA)基准性能已接近饱和。尽管社区不断推出更难数据集,但验证题目质量、筛选错误题目的过程仍高度依赖人力。为此,我们提出一种基于模型输出分布的双组件概率审计框架。首先,在基准层面,通过顶置预测概率(P_{top1})和归一化残差熵(H_{norm})刻画概率景观,以均对距离(MPD)进行全局总结;其次,在题目层面,引入噪声注入以削弱有意义的干扰项竞争,从而识别需人工审查的候选题目,并分类残留错误模式。在四个MCQA基准上的分析显示,不同基准存在模型置信度与残差选项竞争的结构性差异。同时,噪声注入方法所标记的题目与MMLU-Redux专家标注的错误高度一致。结果表明,该概率框架可作为轻量级审计工具,用于比较宏观基准结构并优先筛选个别题目供人工复核。

原文摘要 · Abstract (English)

As Large Language Models rapidly advance, performance on standard multiple-choice question answering (MCQA) benchmarks is reaching saturation. While the community has responded by developing increasingly difficult datasets, validating question quality and filtering flawed items remains a labor-intensive process. To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions. First, for benchmark-level analysis, we characterize the probability landscape using the top prediction probability ($P_{top1}$) and normalized residual entropy ($H_{norm}$), summarized globally by Mean Pairwise Distance (MPD). Second, for item-level diagnostics, we introduce noise injection to reduce meaningful distractor competition, enabling us to flag candidate items for targeted human review and categorize residual failure patterns. Across four MCQA benchmarks, our landscape analysis reveals benchmark-level differences in model confidence and residual option competition. Concurrently, our noise-injection method flags potentially actionable item-level issues, showing alignment with expert error annotations from MMLU-Redux. These results suggest that our probability-based framework provides a lightweight audit lens for comparing macro-level benchmark structure and prioritizing individual items for targeted human review.

模型评估多选题概率分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。