用100个选项测试大模型,发现低选项题会夸大模型能力。
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
- 将选择题选项扩至100个,减少随机猜对的影响。
- 高干扰下模型表现明显下降,暴露原有评估的盲区。
- 适合关注模型真实推理能力的研究者使用。
多选题评估广泛用于大语言模型基准测试,但在选项较少时,模型可能通过捷径策略获得接近满分的成绩,掩盖真实能力。为此,我们提出一种大规模选项评估协议,将候选集扩展至100个选项,显著降低随机性能的影响。该框架应用于韩语拼写纠错任务,要求模型从大量候选句中选出唯一错误句。通过固定目标并反复重采样与打乱,获得稳定评估结果,同时区分内容性失败与位置偏差。实验表明,低选项设置下的强表现常高估模型能力;在高干扰(N=100)条件下,这种优势明显减弱,暴露出传统基准忽略的差距。我们识别出两种失效模式:语义混淆和不确定时对靠前选项的位置偏好。通过控制填充与长度匹配测试,发现主要瓶颈是候选排序而非上下文长度。这些结果支持大规模选项评估作为压力测试模型在极端干扰下可靠性的通用框架。
原文摘要 · Abstract (English)
Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by shortcut strategies that obscure true competence. Therefore, we propose a massive option evaluation protocol that scales the candidate set to one hundred options and sharply reduces the impact of chance performance. We apply this framework to a Korean orthography error detection task where models must pick the single incorrect sentence from a large candidate set. With fixed targets and repeated resampling and shuffling, we obtain stable estimates while separating content driven failures from positional artifacts. Across experiments, results indicate that strong performance in low option settings can overstate model competence. This apparent advantage often weakens under dense interference at high $N$, revealing gaps that conventional benchmarks tend to obscure. We identify two failure modes, semantic confusion and position bias toward early options under uncertainty. To isolate the effect of context length, we run padding controlled and length matched tests, which suggest that the main bottleneck is candidate ranking rather than context length. Together, these findings support massive option evaluation as a general framework for stress testing model reliability under extreme distractor density, beyond what low option benchmarks can reveal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。