研究不同重采样方法在因果发现中的效果,给出实际选择建议。
An extensive simulation study evaluating the interaction of resampling techniques across multiple causal discovery contexts
- 通过理论证明重采样方法等价于调整算法参数
- 大规模模拟验证了理论结果并揭示方法依赖性
- 适合做因果推断的科研人员参考方法选择
尽管探索性因果分析在现代科学和医学中日益普及,但用于验证因果模型的非实验方法尚未得到充分描述。一种常用方法是通过重采样数据评估模型特征的稳定性,类似于统计学中估计置信区间的重采样方法。然而,重采样方法的选择是否应取决于样本量、算法类型或算法调参参数等问题仍缺乏关注。本文提出理论证明,某些重采样方法在数学上等价于对算法调参进行特定赋值。同时报告了大规模模拟实验结果,验证了理论结论,并为研究人员进一步理解因果发现中的重采样提供了丰富数据支持。理论与仿真共同为实践中如何选择重采样方法与调参提供具体指导。
原文摘要 · Abstract (English)
Despite the accelerating presence of exploratory causal analysis in modern science and medicine, the available non-experimental methods for validating causal models are not well characterized. One of the most popular methods is to evaluate the stability of model features after resampling the data, similar to resampling methods for estimating confidence intervals in statistics. Many aspects of this approach have received little to no attention, however, such as whether the choice of resampling method should depend on the sample size, algorithms being used, or algorithm tuning parameters. We present theoretical results proving that certain resampling methods closely emulate the assignment of specific values to algorithm tuning parameters. We also report the results of extensive simulation experiments, which verify the theoretical result and provide substantial data to aid researchers in further characterizing resampling in the context of causal discovery analysis. Together, the theoretical work and simulation results provide specific guidance on how resampling methods and tuning parameters should be selected in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。