arXiv:2606.23880cs.LGstat.ME2026-06被引 1

用门控机制提升高维时间序列因果边发现的准确率。

GRACE: Gated Refinement for Accurate Causal Edge Discovery in High-Dimensional Time Series

  • 引入硬混凝土门控与L0正则化,实现边缘选择的二值化分离。
  • 在d=100的合成数据上F1显著提升,比非线性CI测试快75倍。
  • 适用于存在混淆、时滞等复杂现实场景,适合真实时间序列分析。

从气候遥相关到基因调控,现代时间序列数据包含数十甚至数百个交互变量,导致因果发现愈发困难。基于约束的方法具有统计严谨性,但其非线性条件独立检验难以规模化;而基于评分的方法虽避免了检验,却需人为设定阈值来二值化连续边得分。本文提出GRACE(Gated Refinement for Accurate Causal Edge discovery),通过硬混凝土门控与L0正则化对基于约束的发现进行精炼:每个候选边拥有独立门控,取值集中在0或1附近,实现清晰的双峰分离,使二值决策更鲁棒,优于L1和注意力方法产生的窄且重叠的得分分布。先通过快速线性条件独立骨架获取高召回候选边;再由单一门控模型学习哪些边能真正提升预测能力,自动适应问题维度与骨架密度。在涵盖不同图结构(尺度无标度、Erdős-Rényi、小世界)及维度高达d=100的合成基准上,GRACE显著提升基线条件独立方法的F1,同时保持高精度,并超越注意力与基于评分的替代方案。其性能接近昂贵的非线性条件独立测试,但速度达其75倍。在真实河流流量数据集上,其时序自助变体在埃尔布河中恢复出11条因果边中的9条,仅1个误报(F1=0.86,AUROC=0.99),将骨架的106个误报减少99%。

原文摘要 · Abstract (English)

From climate teleconnections to gene regulation, modern time-series datasets encompass tens or hundreds of interacting variables, making causal discovery increasingly challenging. Constraint-based methods offer statistical rigor but their nonlinear CI tests are infeasible at scale, while score-based alternatives avoid CI testing but require arbitrary thresholds to binarize continuous edge scores. We propose GRACE ($\textbf{G}$ated $\textbf{R}$efinement for $\textbf{A}$ccurate $\textbf{C}$ausal $\textbf{E}$dge discovery), which refines constraint-based discovery using Hard Concrete gates with $L_0$ regularization: each candidate edge has an independent gate whose values concentrate near 0 or 1, yielding a clean bimodal separation that makes the binary decision robust, unlike the narrow, overlapping score distributions produced by $L_1$ and attention-based methods. A fast linear CI skeleton provides high-recall candidates; a single gated model then prunes false positives by learning which edges genuinely improve prediction, with automatic regularization adapted to problem dimensions and skeleton density. Systematic experiments on synthetic benchmarks, spanning diverse graph topologies (scale-free, Erdős-R'enyi, small-world) and dimensionalities up to $d=100$, show that GRACE substantially improves F1 over its base CI method while maintaining high precision, and outperforms attention-based and score-based alternatives. GRACE matches or exceeds expensive nonlinear CI tests at a fraction of the cost ($75\times$ faster). On a real-world river flow dataset, where rainfall confounders, variable propagation lags, and distributional shifts violate standard assumptions, a temporal bootstrap variant of GRACE recovers 9 of 11 causal edges along the Elbe River with only 1 false positive ($F_1 = 0.86$, AUROC${} = 0.99$), reducing the skeleton's 106 false positives by 99%.

因果发现时间序列门控机制高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。