揭示多模态长链推理中知识冲突的四种内在机制。
Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning
- 区分输入层与过程层的知识冲突类型,发现其可线性分离。
- 冲突信号集中在模型中后层,且可通过轨迹聚合恢复冲突类型。
- 模型在冲突中更易强化默认来源偏好,适合调试推理错误。
多模态大语言模型在长链思维推理中常因不同知识源提供矛盾信号而失败。本文提出统一的知识冲突概念,区分输入层客观冲突与过程层有效冲突。通过探测内部表示,发现:(I) 线性可分性:不同冲突类型以线性可分特征显式编码,而非纠缠;(II) 深度定位性:冲突信号集中于中后层,表明存在专门的冲突处理阶段;(III) 层次一致性:沿推理轨迹聚合噪声级的标记信号,可稳健恢复输入层冲突类型;(IV) 方向不对称性:强化模型隐含的来源偏好远比强制相反来源容易。研究为多模态推理中的知识冲突提供了机制层面的理解,并支持系统性诊断与控制长链推理失败。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) in long chain-of-thought reasoning often fail when different knowledge sources provide conflicting signals. We formalize these failures under a unified notion of knowledge conflict, distinguishing input-level objective conflict from process-level effective conflict. Through probing internal representations, we reveal that: (I) Linear Separability: different conflict types are explicitly encoded as linearly separable features rather than entangled; (II) Depth Localization: conflict signals concentrate in mid-to-late layers, indicating a distinct processing stage for conflict encoding; (III) Hierarchical Consistency: aggregating noisy token-level signals along trajectories robustly recovers input-level conflict types; and (IV) Directional Asymmetry: reinforcing the model's implicit source preference under conflict is far easier than enforcing the opposite source. Our findings provide a mechanism-level view of multimodal reasoning under knowledge conflict and enable principled diagnosis and control of long-CoT failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。