发现常识推理题中多数选项都合理,但标准答案常不最合理。
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
- 通过5000次独立判断,评估每个选项的合理性。
- 超20%题目中,最合理答案与标准答案不符。
- 该方法可识别出有歧义或不匹配的低质量题目。
常识推理中的多选题通常有多个可能合理的答案,但现有基准要求选择唯一正确答案,理论上应为最合理的选项。我们从两个常识推理基准中抽取250道题目,收集了5000次独立的选项合理性判断。结果发现,超过20%的题目中,被普遍认为最合理的选项与基准的正确答案不一致;人工检查确认这些题目存在模糊性或问题与选项语义不匹配等问题。对大语言模型的实验显示,这些题目上模型准确率低且表现波动大,表明基于合理性的评估标准有助于筛选更可靠的评测题目。
原文摘要 · Abstract (English)
Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard selection of a single correct answer, which, in principle, should represent the $\textit{most}$ plausible answer choice. On $250$ MCQ items sampled from two commonsense reasoning benchmarks, we collect $5,000$ independent plausibility judgments on answer choices. We find that for over 20% of the sampled MCQs, the answer choice rated most plausible does not match the benchmark gold answers; upon manual inspection, we confirm that this subset exhibits higher rates of problems like ambiguity or semantic mismatch between question and answer choices. Experiments with LLMs reveal low accuracy and high variation in performance on the subset, suggesting our plausibility criterion may be helpful in identifying more reliable benchmark items for commonsense evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。