研究发现大模型在选择题中仅靠选项也能答对,且推理过程能提升准确率。
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
- 让大模型在完整题目和仅选项两种条件下答题,对比推理效果。
- 一半情况下,仅用选项也能提升准确率,且推理长度不影响结果。
- 推理过程揭示模型通过补全问题来作答,非简单捷径,适合评估模型思维质量。
大型语言模型(LLMs)在多选题问答(MCQA)中先进行推理再作答,表现优异。然而,有研究指出,不带推理的模型仅凭选项就能完成任务,可能依赖冗余捷径。为检验此类部分输入的成功是否源于浅层策略,我们让具备推理能力的模型在完整题目与仅选项两种输入下解题。实验发现,测试时推理可提升两类条件下的准确率,其中一半情况在仅选项输入下也有效。尽管可能存在浅层捷径,但仅选项成功几乎不受推理长度影响,且经过忠实性检验后,发现模型实际采用更合理的策略,如推断缺失的问题。因此,我们质疑‘部分输入成功即为缺陷’的普遍结论,并提出利用推理轨迹区分低质量问题与合理推理。
原文摘要 · Abstract (English)
Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern is that LLMs do not solve MCQs as intended, as work finds LLMs sans reasoning succeed in MCQA without using the question, i.e., choices-only. Such partial-input success is often linked to trivial shortcuts, but reasoning traces could reveal if choices-only strategies are truly shallow. To examine these strategies, we have reasoning LLMs solve MCQs in full and choices-only inputs; test-time reasoning often boosts accuracy in full and in choices-only, half the time. While possibly due to shallow shortcuts, choices-only success is barely affected by the length of reasoning traces, and after finding traces pass faithfulness tests, we show they use less problematic strategies like inferring missing questions. In all, we challenge claims that partial-input success is always a flaw, so we propose how reasoning traces could separate problematic data from less problematic reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。