测试时间序列因果发现方法在假设失效下的鲁棒性,揭示33种失效情形的表现差异。
TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations

- 构建可定制的TCD-Arena测试工具,逐步引入假设破坏
- 完成约3000万次实验,分析33种假设失效的鲁棒性表现
- 发现集成方法能提升整体鲁棒性,适合实际应用研究者
因果发现(CD)是科学探索的强大框架,但其实际应用受限于对强且常不可验证假设的依赖,以及缺乏稳健性能评估。为解决这些问题并推进实证评估,我们提出TCD-Arena,一个模块化、高度可定制且可扩展的测试工具,用于评估时间序列因果发现算法在逐步加剧的假设违反情况下的鲁棒性。作为演示,我们进行了涵盖约3000万次独立因果发现尝试的广泛实证研究,揭示了33种不同假设违反情形下的细微鲁棒性特征。此外,我们研究了因果发现集成方法,发现其具备提升整体鲁棒性的潜力,这对真实世界应用具有重要意义。我们的目标是最终推动开发出在多样合成数据及潜在真实数据条件下均可靠的因果发现方法。
原文摘要 · Abstract (English)
Causal Discovery (CD) is a powerful framework for scientific inquiry. Yet, its practical adoption is hindered by a reliance on strong, often unverifiable assumptions and a lack of robust performance assessment. To address these limitations and advance empirical CD evaluation, we present TCD-Arena, a modularized, highly customizable, and extendable testing kit to assess the robustness of time series CD algorithms against stepwise more severe assumption violations. For demonstration, we conduct an extensive empirical study comprising around 30 million individual CD attempts and reveal nuanced robustness profiles for 33 distinct assumption violations. Further, we investigate CD ensembles and find that they have the potential to improve general robustness, which has implications for real-world applications. With this, we strive to ultimately facilitate the development of CD methods that are reliable for a diverse range of synthetic and potentially real-world data conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。