arXiv:2608.06938cs.CVcs.AI2026-08

用文本引导模型克服先验偏见,提升视觉反常识推理能力

Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning

论文配图:Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
图 1 · 摘自论文原文
  • 通过文本构建反常识场景数据集,强化低频事实的语义表征
  • 在多个视觉反常识基准上性能显著提升,准确率最高达78.3%
  • 无需视觉数据微调,适合需增强推理公平性的多模态应用

多模态大语言模型(MLLMs)的视觉推理能力对下游任务至关重要,尤其在反常识推理中,需突破普遍认知假设。现有研究多聚焦于增强视觉输入,认为失败源于视觉基础不足。然而实证分析表明,瓶颈并非视觉感知:MLLMs 已能捕捉相关视觉证据,正确答案存在于其解码空间。问题在于共享语言解码器在先验与证据冲突时偏好主流语言先验,尤其在低频事实场景中。为此,我们提出文本锚定的数据构建流程,核心为事实频率蒸馏(FFD),用于估计常识事实的先验强度,并将验证过的反常识场景提炼为高质量文本语料库。基于此语料库,我们提出 TACT 框架——一种无需视觉训练数据的文本锚定后训练方法。TACT 将证据导向与先验驱动的推理路径分阶段优化,使解码器有效化解先验-证据冲突。在多个反常识视觉基准测试中,TACT 显著提升视觉推理性能,同时保持通用能力,验证了有效的文本到视觉跨模态迁移。

原文摘要 · Abstract (English)

The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions. Recent studies mainly improve visual counter-commonsense reasoning by enhancing visual inputs, following the assumption that failures originate from insufficient visual grounding. However, our empirical analysis reveals that the bottleneck is not visual perception. MLLMs already capture the relevant visual evidence, and the correct answer exists in their decoding space. Instead, the shared language decoder resolves prior--evidence conflicts by favoring dominant language priors, especially for low-frequency factual scenarios. Motivated by this, we first propose a text-anchored data construction pipeline, whose core component, Fact-Frequency Distillation (FFD), estimates the prior strength of commonsense facts and distills verified counter-commonsense scenarios into a high-quality text corpus. Building upon this corpus, we introduce TACT, a text-anchored post-training framework that debiases the shared language decoder without requiring any visual training data. TACT routes evidence-following and prior-driven reasoning trajectories into different optimization stages, enabling the decoder to resolve prior--evidence conflicts. Across counter-commonsense visual benchmarks, TACT substantially improves visual reasoning while preserving general capabilities, demonstrating effective text-to-vision cross-modal transfer.

多模态推理反常识推理去偏文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。