推理模型更会解释自己为何作答,可信度显著提升。
Are DeepSeek R1 And Other Reasoning Models More Faithful?
- 用强化学习训练的推理模型更擅长描述提示中线索如何影响答案
- DeepSeek-R1能解释线索影响的比例达59%,远超非推理模型的7%
- 适合关注模型可解释性与决策透明度的研究者和开发者
通过强化学习训练的推理语言模型在解决推理任务上表现优异。我们评估了基于Qwen-2.5、Gemini-2和DeepSeek-V3-Base的三类推理模型,在一项检验思维链(CoT)忠实性的测试中表现。通过七种类型线索(如误导性示例、用户引导问题)考察模型是否能准确描述提示中线索如何影响其对MMLU题目的回答。例如,当提示加入“一位斯坦福教授认为答案是D”时,推理模型在59%情况下能说明该线索对其判断的影响,而对应非推理模型仅7%。所有推理模型均显著优于非推理模型(包括Claude-3.5-Sonnet和GPT-4o)。额外实验表明,奖励模型的使用可能降低响应忠实性,或解释了非推理模型表现不佳的原因。研究局限在于测试基于人工任务,且仅衡量线索影响的可解释性。未来需拓展至更广泛场景。但当前结果为提升语言模型可解释性提供了积极信号。
原文摘要 · Abstract (English)
Language models trained to solve reasoning tasks via reinforcement learning have achieved striking results. We refer to these models as reasoning models. Are the Chains of Thought (CoTs) of reasoning models more faithful than traditional models? We evaluate three reasoning models (based on Qwen-2.5, Gemini-2, and DeepSeek-V3-Base) on an existing test of faithful CoT. To measure faithfulness, we test whether models can describe how a cue in their prompt influences their answer to MMLU questions. For example, when the cue "A Stanford Professor thinks the answer is D" is added to the prompt, models sometimes switch their answer to D. In such cases, the DeepSeek-R1 reasoning model describes the cue's influence 59% of the time, compared to 7% for the non-reasoning DeepSeek model. We evaluate seven types of cue, such as misleading few-shot examples and suggestive follow-up questions from the user. Reasoning models describe cues that influence them much more reliably than all the non-reasoning models tested (including Claude-3.5-Sonnet and GPT-4o). In an additional experiment, we provide evidence suggesting that the use of reward models causes less faithful responses -- which may help explain why non-reasoning models are less faithful. Our study has two main limitations. First, we test faithfulness using a set of artificial tasks, which may not reflect realistic use-cases. Second, we only measure one specific aspect of faithfulness -- whether models can describe the influence of cues. Future research should investigate whether the advantage of reasoning models in faithfulness holds for a broader set of tests. Still, we think this increase in faithfulness is promising for the explainability of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。