arXiv:2502.12425cs.CV2025-02TPAMI被引 23

通过解耦与反事实学习,提升模型在音视频缺失时的物理常识推理能力。

Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

  • 用变分自编码器解耦视频的静态与动态特征,增强多模态表征
  • 引入反事实模块模拟物体间物理关系,提升因果推理能力
  • 可插拔设计,适用于各类视觉语言模型,尤其适合模态缺失场景

本文提出一种鲁棒解耦反事实学习(RDCL)方法,用于物理音视频常识推理。该任务基于视频与音频输入推断物体的物理常识,核心挑战在于如何模拟人类在模态缺失情况下的推理能力。现有方法未能充分利用多模态数据差异,且模型缺乏因果推理能力,阻碍了隐式物理知识的推断。为此,本方法通过解耦序列编码器将视频在潜在空间中分解为静态(时间不变)和动态(时间可变)因素,采用变分自编码器(VAE)结合对比损失函数以最大化互信息。进一步引入反事实学习模块,通过模拟反事实干预下不同物体间的物理关系来增强模型推理能力。为缓解模态缺失问题,提出鲁棒多模态学习方法,通过分离共享特征与模态特有特征实现缺失数据恢复。所提方法为即插即用模块,可集成于任意基线模型,包括视觉语言模型(VLMs)。实验表明,该方法显著提升基线模型的推理准确率与鲁棒性,达到当前最优性能。

原文摘要 · Abstract (English)

In this paper, we propose a new Robust Disentangled Counterfactual Learning (RDCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects' physics commonsense based on both video and audio input, with the main challenge being how to imitate the reasoning ability of humans, even under the scenario of missing modalities. Most of the current methods fail to take full advantage of different characteristics in multi-modal data, and lacking causal reasoning ability in models impedes the progress of implicit physical knowledge inferring. To address these issues, our proposed RDCL method decouples videos into static (time-invariant) and dynamic (time-varying) factors in the latent space by the disentangled sequential encoder, which adopts a variational autoencoder (VAE) to maximize the mutual information with a contrastive loss function. Furthermore, we introduce a counterfactual learning module to augment the model's reasoning ability by modeling physical knowledge relationships among different objects under counterfactual intervention. To alleviate the incomplete modality data issue, we introduce a robust multimodal learning method to recover the missing data by decomposing the shared features and model-specific features. Our proposed method is a plug-and-play module that can be incorporated into any baseline including VLMs. In experiments, we show that our proposed method improves the reasoning accuracy and robustness of baseline methods and achieves the state-of-the-art performance.

多模态推理反事实学习物理常识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。