评估因果发现结果的可靠性,强调效果验证比图结构准确更重要。
Effect-Level Validation for Causal Discovery
- 以可识别性为先,通过可识别性、稳定性与可证伪性检验因果图假设。
- 多数看似合理的因果图在加入时间语义约束后无法得出点估计,可识别性成瓶颈。
- 不同算法生成不同图结构但效果估计一致,适合用于决策支持系统。
因果发现被广泛应用于大规模遥测数据以评估用户干预的影响,但在强自选择的反馈驱动系统中其决策可靠性尚不明确。本文提出一种以效果为中心、可识别性优先的框架,将发现的因果图视为结构性假设,并通过可识别性、稳定性和可证伪性进行评估,而非仅依赖图恢复准确率。基于真实游戏遥测数据,研究了早期接触竞技玩法对短期留存的影响。结果发现,许多统计上合理的发现输出在施加最小的时间与语义约束后,无法实现点识别的因果查询,凸显可识别性是决策支持的关键瓶颈。当可识别性成立时,多种算法家族虽生成差异显著的图结构,却得出相似且决策一致的效果估计,包括直接处理-结果边缺失而效果通过间接路径保留的情况。这些一致估计经受住安慰剂、子采样和敏感性检验。相反,其他方法表现出零散的可识别性及阈值敏感或衰减的效果,源于终点定义模糊。结果表明,图级指标无法充分代表特定目标查询的因果可靠性。因此,在遥测驱动系统中,可信因果结论需优先保障可识别性与效果级验证,而非仅关注因果结构恢复。
原文摘要 · Abstract (English)
Causal discovery is increasingly applied to large-scale telemetry data to estimate the effects of user-facing interventions, yet its reliability for decision-making in feedback-driven systems with strong self-selection remains unclear. In this paper, we propose an effect-centric, admissibility-first framework that treats discovered graphs as structural hypotheses and evaluates them by identifiability, stability, and falsification rather than by graph recovery accuracy alone. Empirically, we study the effect of early exposure to competitive gameplay on short-term retention using real-world game telemetry. We find that many statistically plausible discovery outputs do not admit point-identified causal queries once minimal temporal and semantic constraints are enforced, highlighting identifiability as a critical bottleneck for decision support. When identification is possible, several algorithm families converge to similar, decision-consistent effect estimates despite producing substantially different graph structures, including cases where the direct treatment-outcome edge is absent and the effect is preserved through indirect causal pathways. These converging estimates survive placebo, subsampling, and sensitivity refutation. In contrast, other methods exhibit sporadic admissibility and threshold-sensitive or attenuated effects due to endpoint ambiguity. These results suggest that graph-level metrics alone are inadequate proxies for causal reliability for a given target query. Therefore, trustworthy causal conclusions in telemetry-driven systems require prioritizing admissibility and effect-level validation over causal structural recovery alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。