arXiv:2602.23800stat.MEcs.AI2026-02

将实际工作流程约束融入纵向因果发现,提升模型可解释性与时间一致性。

Operationalizing Longitudinal Causal Discovery Under Real-World Workflow Constraints

  • 基于工作流程生成结构约束,限制因果图搜索空间。
  • 在10万+人群数据中实现时间一致的因果结构与带不确定性的滞后效应分析。
  • 适合需要可复现因果推断的医疗、公共健康研究者。

因果发现虽理论进展显著,但在大规模纵向系统中的应用仍受限。主要障碍在于操作数据受机构工作流程影响,其隐含的偏序关系未被形式化,导致可接受的因果图空间扩大,与记录过程不一致。本文提出一种由工作流程导出的约束类别,通过协议生成的结构掩码和时间对齐索引,限制有向无环图空间。无需新优化算法,显式编码工作流程一致的偏序关系可减少结构歧义,尤其在混合离散-连续面板中效果显著。该框架整合了工作流程衍生的可接受边约束、测量对齐的时间索引、块结构、基于自助法的滞后总效应不确定性量化,以及支持干预查询的动态表示。在日本全国年度健康筛查队列(107,261人,429,044人年)中,工作流程约束的纵向LiNGAM实现了时间一致的组内结构和可解释的滞后效应,并明确给出不确定性。使用不同暴露和体成分定义的敏感性分析保持了主要定性模式。本文认为,形式化工作流程导出的约束类可提升结构可解释性,无需领域特定边设定,在标准可识别性假设下为操作流程与纵向因果发现之间提供可复现的桥梁。

原文摘要 · Abstract (English)

Causal discovery has achieved substantial theoretical progress, yet its deployment in large-scale longitudinal systems remains limited. A key obstacle is that operational data are generated under institutional workflows whose induced partial orders are rarely formalized, enlarging the admissible graph space in ways inconsistent with the recording process. We characterize a workflow-induced constraint class for longitudinal causal discovery that restricts the admissible directed acyclic graph space through protocol-derived structural masks and timeline-aligned indexing. Rather than introducing a new optimization algorithm, we show that explicitly encoding workflow-consistent partial orders reduces structural ambiguity, especially in mixed discrete--continuous panels where within-time orientation is weakly identified. The framework combines workflow-derived admissible-edge constraints, measurement-aligned time indexing and block structure, bootstrap-based uncertainty quantification for lagged total effects, and a dynamic representation supporting intervention queries. In a nationwide annual health screening cohort in Japan with 107,261 individuals and 429,044 person-years, workflow-constrained longitudinal LiNGAM yields temporally consistent within-time substructures and interpretable lagged total effects with explicit uncertainty. Sensitivity analyses using alternative exposure and body-composition definitions preserve the main qualitative patterns. We argue that formalizing workflow-derived constraint classes improves structural interpretability without relying on domain-specific edge specification, providing a reproducible bridge between operational workflows and longitudinal causal discovery under standard identifiability assumptions.

因果发现纵向分析医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。