arXiv:2609.07874cs.LGstat.ML2026-09

提出可干预的因果场模型,让多模态模型更好理解视觉操作的影响。

InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling

论文配图:InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling
图 1 · 摘自论文原文
  • 在视觉编码器与语言解码器间引入可微因果场,实现跨步影响传播。
  • 在CausalVQA上提升13.1个百分点,尤其在规划和假设类任务中表现突出。
  • 适用于需要因果推理的多模态任务,如视觉问答与决策建模。

多模态大语言模型常捕捉视觉-语言关联,但难以预测局部视觉干预如何传播并影响下游回答。本文提出InfluenceField,一个插入在视觉编码器与语言解码器之间的干预感知潜在场。该场将图像块特征升维为连续空间表示,通过共享转移算子实现多步定向影响传播,并预测局部干预效应。联合训练优化语言建模、跨环境不变性、反事实回溯监督与结构正则化。对于非线性有限基群体模型,我们证明目标对齐的干预监督结合一步分离条件,可将可接受表示限制在局部重参数化范围内,从而精确恢复完整转移的有向依赖图。线性特化给出精确部分覆盖刻画与有限损失稳定性界,并推导出系数干预的空间分布及共享通道校准结果。在CausalVQA上,InfluenceField相比基线模型整体准确率提升13.1个百分点,尤其在规划与假设类别中增益最大。容量匹配的基线与结构控制实验表明,鲁棒性与事实-反事实一致性提升源于因果目标而非额外容量。

原文摘要 · Abstract (English)

Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. It lifts patch features into a continuous spatial representation, propagates directed influence over multiple steps, and predicts local intervention effects through a shared transition operator. Training jointly optimizes language modeling, cross-environment invariance, counterfactual rollout supervision, and structural regularization. For a nonlinear finite-basis population model, we show that target-aligned interventional supervision, together with a one-step separation condition on the transition, restricts admissible representations to within-location reparameterizations, so that the directed dependency graph of the full transition is recovered exactly. A linear specialization gives an exact partial-coverage characterization and a finite-loss stability bound, and the field analysis derives the spatial profile of coefficient interventions together with a shared-channel calibration result. On CausalVQA, InfluenceField improves overall accuracy over its backbone by 13.1 percentage points, with the largest gains on the planning and hypothetical categories. Capacity-matched baselines and structural controls attribute the gains in robustness and factual-counterfactual consistency to the causal objectives rather than to added capacity.

因果建模多模态视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。