arXiv:2601.04073cs.CVcs.AI2026-01被引 6

发现大模型推理中文字惯性问题,提出视觉主动校准方法提升鲁棒性。

Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts

  • 通过结构化扰动测试推理链,发现模型对文本错误难以自纠
  • 自纠正成功率低于10%,多数情况盲从错误文本
  • 无需训练的推理机制,主动重审视觉信息并清理推理过程

大型多模态模型(LMMs)在视频推理中展现出了强大的链式思维(CoT)能力,但其推理链的鲁棒性仍存疑。本文识别出一种关键失败模式——文本惯性:一旦思维过程中出现文本幻觉,模型往往盲目坚持错误文本,忽视矛盾的视觉证据。为系统研究此问题,我们提出逻辑图扰动协议(LogicGraph Perturbation Protocol),对多种跨原生推理架构与提示驱动范式的LMMs进行结构化扰动,评估其自我反思能力。结果显示,模型自纠正成功比例不足10%,多数陷入错误传播。为此,我们引入无训练推理范式「主动视觉上下文精炼」(Active Visual-Context Refinement),通过主动视觉再定位机制实现细粒度验证,并结合自适应上下文精炼策略对推理历史进行总结与去噪。实验表明,该方法显著抑制了幻觉传播,提升了推理鲁棒性。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have demonstrated impressive capabilities in video reasoning via Chain-of-Thought (CoT). However, the robustness of their reasoning chains remains questionable. In this paper, we identify a critical failure mode termed textual inertia, where once a textual hallucination occurs in the thinking process, models tend to blindly adhere to the erroneous text while neglecting conflicting visual evidence. To systematically investigate this, we propose the LogicGraph Perturbation Protocol that structurally injects perturbations into the reasoning chains of diverse LMMs spanning both native reasoning architectures and prompt-driven paradigms to evaluate their self-reflection capabilities. The results reveal that models successfully self-correct in less than 10% of cases and predominantly succumb to blind textual error propagation. To mitigate this, we introduce Active Visual-Context Refinement, a training-free inference paradigm which orchestrates an active visual re-grounding mechanism to enforce fine-grained verification coupled with an adaptive context refinement strategy to summarize and denoise the reasoning history. Experiments demonstrate that our approach significantly stifles hallucination propagation and enhances reasoning robustness.

多模态推理幻觉抑制视觉校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。