arXiv:2505.20214cs.AI2025-05ACL

慢推理模型在多模态任务中反而更易说谎,因陷入错误前提深挖。

When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning

  • 用分层提示数据集对比不同推理模式,发现慢速模型倾向深度优先探索错误前提。
  • 在5000个含误导视觉输入的样本中,慢推理模型虚假细节生成率显著更高。
  • 适合关注大模型幻觉机制、多模态推理可信度的研究者阅读。

推理模型因其处理复杂任务的能力而备受关注,体现了与快速直觉反应(系统I)相对的系统II(慢思考)范式。然而一个关键问题仍待解答:更慢的推理是否必然带来更真实的回答?我们的研究发现并非如此。我们首次系统性地探究了多模态推理中慢思考范式的反向缩放规律。当面对不完整或具有误导性的视觉输入时,慢思考模型更倾向于编造看似合理却虚假的细节来支撑错误推理。为分析此行为,我们构建了一个包含5000个样本的分层提示数据集,由50名人类参与者标注。提示难度逐级递增,揭示出一致规律:慢推理模型倾向于采用深度优先搜索(DFS)思维,持续深挖错误前提;而快速聊天模型则偏好广度优先搜索(BFS)推理,在不确定情况下表现更谨慎。这些发现揭示了推理模型的关键脆弱性:尽管在数学等结构化领域有效,其基于DFS的推理在面对模糊、多模态输入时极易失效。

原文摘要 · Abstract (English)

Reasoning models have attracted increasing attention for their ability to tackle complex tasks, embodying the System II (slow thinking) paradigm in contrast to System I (fast, intuitive responses). Yet a key question remains: Does slower reasoning necessarily lead to more truthful answers? Our findings suggest otherwise. We conduct the first systematic study of the inverse scaling law in slow-thinking paradigms for multimodal reasoning. We find that when confronted with incomplete or misleading visual inputs, slow-thinking models are more prone to fabricating plausible yet false details to justify untruthful reasoning. To analyze this behavior, we construct a 5,000-sample hierarchical prompt dataset annotated by 50 human participants. The prompts progressively increase in complexity, revealing a consistent pattern: slower reasoning models tend to follow depth-first search (DFS) thinking, persistently exploring flawed premises, while faster chat models favor breadth-first search (BFS) inference, showing greater caution under uncertainty. These findings reveal a critical vulnerability of reasoning models: while effective in structured domains such as math, their DFS-style reasoning becomes fragile when confronted with ambiguous, multimodal inputs.

多模态推理幻觉机制推理模型系统二

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。