模型会因错误来源提示而改变正确答案,暴露其判断不稳问题。
When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA
- 用固定错误选项测试不同提示模板下的答案变化
- 专家模板下41.1%的答题不稳定率,远高于多数模板的12.5%
- 即使正确提示准确率高,错误来源提示仍能误导模型
语言模型在回答多选题时,常会收到关于其他来源答案的陈述。本文审计此类陈述是否导致答案不稳定。针对每个题目,固定一个错误选项,在不同误导性提示模板下观察模型选择变化。提出‘中性条件下误导提示采纳率’(NC-MCAR),衡量同一模型在中性提示下本选正确答案,但在有效提示下却转向该错误选项的比例。这反映的是答案稳定性问题,而非模型是否知道正确答案或所有从众行为都非理性。在MMLU-Pro和IndicMMLU-Pro(英语、印地语、孟加拉语、泰米尔语、泰卢固语)上评估四款指令遵循模型,共生成22万条输出。专家模板的平均NC-MCAR达41.1%,而多数模板为12.5%,二者使用相同错误选项与最终指令。填空项准确率显著高于专家错误项准确率,正确提示条件下的有效响应准确率也较高。结果表明,在强制选择提示下,未经验证的来源声明可能压倒原本与任务证据一致的答案。
原文摘要 · Abstract (English)
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\% aggregate NC-MCAR, compared with 12.5\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。