arXiv:2509.19112cs.LGcs.AI2025-09中稿 · NeurIPS被引 3

用一次聚合方法高效发现高维事件序列中的多标签因果关系。

Towards Practical Multi-label Causal Discovery in High-Dimensional Event Sequences via One-Shot Graph Aggregation

  • 通过预训练因果Transformer并行生成每条序列的因果图。
  • 在包含29,100种事件类型的数据上实现474个标签的准确因果推断。
  • 适合医疗、车载诊断等高维稀疏事件场景的可扩展因果分析。

理解疾病或系统故障等结果标签由症状或错误码等前序事件引发的因果关系至关重要,但仍是医疗与车辆诊断等领域未解难题。我们提出CARGO,一种针对稀疏高维事件序列(含数千种唯一事件类型)的可扩展多标签因果发现方法。利用两个预训练因果Transformer作为领域特定基础模型,CARGO并行地为每条序列一次性生成因果图,并通过自适应频率融合进行聚合,重构标签的全局马尔可夫边界。该两阶段方法在避免全数据集条件独立性测试带来的不可行成本的同时,实现大规模概率推理。在包含超过29,100种唯一事件类型和474个不平衡标签的真实汽车故障预测数据集上,实验验证了CARGO具备结构化推理能力。

原文摘要 · Abstract (English)

Understanding causality in event sequences where outcome labels such as diseases or system failures arise from preceding events like symptoms or error codes is critical. Yet remains an unsolved challenge across domains like healthcare or vehicle diagnostics. We introduce CARGO, a scalable multi-label causal discovery method for sparse, high-dimensional event sequences comprising of thousands of unique event types. Using two pretrained causal Transformers as domain-specific foundation models for event sequences. CARGO infers in parallel, per sequence one-shot causal graphs and aggregates them using an adaptive frequency fusion to reconstruct the global Markov boundaries of labels. This two-stage approach enables efficient probabilistic reasoning at scale while bypassing the intractable cost of full-dataset conditional independence testing. Our results on a challenging real-world automotive fault prediction dataset with over 29,100 unique event types and 474 imbalanced labels demonstrate CARGO's ability to perform structured reasoning.

因果发现事件序列多标签高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。