arXiv:2606.13024cs.LGcs.AI2026-06KDD

百亿参数模型动态识别时间模式,精准发现复杂系统中的因果关系。

CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

论文配图:CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts
图 1 · 摘自论文原文
  • 用模式路由异质专家,按时间片段分配专用模型处理
  • 在监督和少样本场景下均超越现有方法,实现稀疏可解释因果图
  • 首次融合大语言和视觉模型,用文本图像先验约束因果推断

格兰杰因果发现(GCD)是分析复杂系统中时间依赖性的基础。然而,现有神经GCD方法多采用“一刀切”范式,难以捕捉真实时间序列中的分布变化和动态制度转换,常导致表征纠缠和虚假因果图。本文提出CausalMoE,一个百亿级多模态格兰杰因果基础模型,显式建模片段级异质性。该模型引入模式路由异质专家混合机制,动态识别潜在时间模式,并将时间片段路由至专业领域专家,有效分离特定制度机制与共享动态。为确保可解释的图恢复,设计了跨变量的因果感知自注意力机制,通过近端优化生成稀疏格兰杰因果图。此外,CausalMoE首次整合大语言模型(LLM)与视觉语言模型(VLM),将数值信号与文本、视觉先验对齐,在复杂场景中正则化因果估计。大量实验表明,CausalMoE在全监督基准上建立新SOTA,且在传统方法失效的少样本设置中仍具强泛化能力。

原文摘要 · Abstract (English)

Granger Causal Discovery (GCD) is fundamental for analyzing temporal dependencies in complex systems. However, existing neural GCD methods predominantly rely on a "one-size-fits-all" paradigm, struggling to capture distribution shifts and dynamic regime changes inherent in real-world time series. This often leads to entangled representations and spurious causal graphs. In this paper, we propose CausalMoE, a billion-scale multimodal Granger causal foundation model that explicitly models patch-level heterogeneity. CausalMoE introduces a Pattern-Routed Mixture of Heterogeneous Experts, which dynamically identifies latent temporal patterns and routes patches to specialized domain experts, effectively decoupling regime-specific mechanisms from shared dynamics. To ensure interpretable graph recovery, we design a Causality-Aware Self-Attention mechanism operating across variables, yielding sparse Granger causal graphs via proximal optimization. Furthermore, CausalMoE is the first to integrate LLMs and VLMs to align numerical signals with textual and visual priors, regularizing causal estimation in complex scenarios. Extensive experiments demonstrate that CausalMoE establishes a new state-of-the-art on fully supervised benchmarks, while effectively generalizing to few-shot settings where traditional methods fail.

因果发现多模态大模型异质专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。