arXiv:2507.06603cs.CV2025-07被引 2

通过跨模态因果干预,提升长时动作识别的准确性与鲁棒性。

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

  • 构建视频与文本间的结构化因果模型,揭示跨模态因果关系。
  • 在Charades、Breakfast、COIN三数据集上性能优于现有方法。
  • 适合研究多模态学习与因果推理的科研人员参考。

长时动作识别(LTAR)因时间跨度长、原子动作关联复杂及视觉混淆因素而具挑战性。尽管视觉-语言模型(VLMs)展现出潜力,但通常依赖统计相关性而非因果机制。现有基于因果的方法仅处理单模态偏差,缺乏跨模态因果建模,限制了其在基于VLM的LTAR中的应用。本文提出跨模态双因果学习(CMDCL),引入结构化因果模型以挖掘视频与标签文本间的因果关系。通过文本因果干预消除文本嵌入中的跨模态偏差,并利用去偏文本引导视觉因果干预,去除视觉模态固有的混淆因子。双重因果干预使模型生成更鲁棒的动作表征,有效应对LTAR挑战。在Charades、Breakfast和COIN三个基准数据集上的实验结果验证了该方法的有效性。代码已公开于https://github.com/xushaowu/CMDCL。

原文摘要 · Abstract (English)

Long-term action recognition (LTAR) is challenging due to extended temporal spans with complex atomic action correlations and visual confounders. Although vision-language models (VLMs) have shown promise, they often rely on statistical correlations instead of causal mechanisms. Moreover, existing causality-based methods address modal-specific biases but lack cross-modal causal modeling, limiting their utility in VLM-based LTAR. This paper proposes \textbf{C}ross-\textbf{M}odal \textbf{D}ual-\textbf{C}ausal \textbf{L}earning (CMDCL), which introduces a structural causal model to uncover causal relationships between videos and label texts. CMDCL addresses cross-modal biases in text embeddings via textual causal intervention and removes confounders inherent in the visual modality through visual causal intervention guided by the debiased text. These dual-causal interventions enable robust action representations to address LTAR challenges. Experimental results on three benchmarks including Charades, Breakfast and COIN, demonstrate the effectiveness of the proposed model. Our code is available at https://github.com/xushaowu/CMDCL.

动作识别因果学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。