arXiv:2608.15428cs.CLcs.AI2026-08

模型可能靠选项位置猜题,而非理解问题。

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

  • 用选项顺序测试模型是否依赖位置盲选
  • 8个选项顺序下11.8%题目全答对,远超随机期望
  • 换题干或改干扰项仍可被破解,说明格式漏洞

多选题评测只看答案选项是否正确,不检验模型是否理解问题。我们在乌克兰法官资格委员会发布的UA-JudgeExam数据集上验证:该数据集含11,990道四选项题,有官方答案。当仅给选项不给问题时,Claude Haiku 4.5得分0.383(高于随机0.25),其中11.8%的题目在所有八种选项顺序下均答对,远超随机预期的0.2题。检索28万份乌克兰法律文本后,匹配度仅0.128。剔除可被检索出的题目后,剩余8,128题中,门控模型自身得分0.204;而未参与筛选的GPT-5.6在无问题情况下仍能答对0.515。对12个模型逐一扣除其选项偏好后,仅GPT-5.6(+0.265)和Sonnet 4.6(+0.081)仍有显著优势。若不剔除偏好,排名严重失真:Llama 3.1 8B因92%题目选A,盲答得分0.292,高于多数模型。门控机制确能识别真实难题,在拒绝的题目上,11个模型得分0.518–0.789,显著高于保留题目的表现。但该信号仅属单模型,无法推广。在400题采样中,9个模型表现与随机无异。重写干扰项后得分降至0.168,低于随机且仍可被利用。在LEXam数据集上,所有选项长度均小于33字符,且指向题干,探测显示无漏洞。题型格式决定漏洞是否存在,模型能力决定漏洞可被提取程度。论文公开了数据集、预测结果和评估工具包。

原文摘要 · Abstract (English)

Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.

测评漏洞多选题模型偏见法律AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。