优化选择题干扰项设计,让大模型更深入推理而非猜答案
Rethinking Multiple-Choice Questions for RLVR: Unlocking Potential via Distractor Design
- 通过迭代构建高质量干扰项,阻止模型简单排除法作弊
- 2选1题也能有效训练,只要干扰项足够强
- 适合想提升大模型逻辑推理能力的研究者
强化学习结合可验证奖励(RLVR)显著提升了大语言模型的推理能力。在应用中,多选题(MCQs)提供了可扩展的可验证数据源,但可能引发奖励劫持——模型通过随机猜测或简单排除法绕过真实推理。现有方法常将多选题转为开放题以规避问题,却损失了专家设计干扰项带来的对比信号。本文系统研究选项设计对RLVR的影响,发现:(1) 训练与测试时选项数量不一致会降低性能;(2) 强有力的干扰项能有效抑制随机猜测,使2选1题目仍可用于高效RLVR训练。基于此,我们提出迭代干扰项精炼(IDC)框架,主动构建高质量干扰项以阻断排除法捷径,促进深层推理。在多个基准上的实验表明,该方法显著提升干扰项质量,并在RLVR训练中带来明显性能增益。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capabilities of Large Language Models. When applied to RLVR, Multiple-Choice Questions (MCQs) offer a scalable source of verifiable data but risk inducing reward hacking, where models shortcut reasoning via random guessing or simple elimination. Current approaches often mitigate this by converting MCQs to open-ended formats, thereby discarding the contrastive signal provided by expert-designed distractors. In this work, we systematically investigate the impact of option design on RLVR. Our analysis highlights two primary insights: (1) Mismatches in option counts between training and testing degrade performance. (2) Strong distractors effectively mitigate random guessing, enabling effective RLVR training even with 2-way questions. Motivated by these findings, we propose Iterative Distractor Curation (IDC), a framework that actively constructs high-quality distractors to block elimination shortcuts and promote deep reasoning. Experiments on various benchmarks demonstrate that our method effectively enhances distractor quality and yields significant gains in RLVR training compared to the original data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。