测试大模型在无正确选项时能否拒绝回答,发现对齐反而让模型更易选错。
Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
- 设计新评测框架,检验模型面对无效选项时的拒绝能力。
- 对齐模型常盲目选错,基础模型拒绝率随规模增大而提升。
- 揭示对齐训练可能削弱模型批判性判断,适合关注AI鲁棒性的研究者。
本文提出一种新框架,评估大语言模型在多项选择题中无有效答案时,平衡指令遵循与批判性推理的能力。通过在算术、领域知识及高风险医疗决策任务上的系统评估,我们发现后训练对齐模型往往默认选择无效选项,而基础模型展现出随规模增长的拒绝能力。分析表明,尽管对齐旨在提升帮助性,却可能意外损害模型的反思判断力——即在面对无效选项时主动拒绝的能力。我们还进行了平行人类研究,发现人类也存在类似指令遵循偏差,提示此类偏差可能通过人类反馈数据集传播至对齐过程。通过广泛的消融实验,考察了模型规模、训练方法与提示工程的影响。研究揭示了对齐优化与批判性推理能力保留之间的根本矛盾,对构建真实场景下更稳健的AI系统具有重要启示。
原文摘要 · Abstract (English)
This work introduces a novel framework for evaluating LLMs' capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. Through systematic evaluation across arithmetic, domain-specific knowledge, and high-stakes medical decision tasks, we demonstrate that post-training aligned models often default to selecting invalid options, while base models exhibit improved refusal capabilities that scale with model size. Our analysis reveals that alignment techniques, though intended to enhance helpfulness, can inadvertently impair models' reflective judgment--the ability to override default behaviors when faced with invalid options. We additionally conduct a parallel human study showing similar instruction-following biases, with implications for how these biases may propagate through human feedback datasets used in alignment. We provide extensive ablation studies examining the impact of model size, training techniques, and prompt engineering. Our findings highlight fundamental tensions between alignment optimization and preservation of critical reasoning capabilities, with important implications for developing more robust AI systems for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。