用生物约束提升碎片化医学数据的因果建模能力
RetiSEM: Generalising Causal Models for Fragmented Biomedical Data

- 将变量按生物学意义分组,加入禁止边约束,分解路径效应
- 合成数据上结构误差更低,真实数据中视网膜指标为下游标志物
- 适合资源有限但需可解释因果推断的医学研究者使用
从碎片化的生物医学数据中学习因果模型极具挑战,因临床、分子与影像变量常不完整或未联合观测。本文提出RetiSEM,一种领域约束的结构方程模型框架,用于在多模态资源有限下进行因果图恢复与中介分析。该方法将变量组织为生物学合理的区块,施加禁止边约束,并将路径级效应分解为总效应(TE)、自然直接效应(NDE)与自然间接效应(NIE)。我们在十种合成基准场景中评估,涵盖维度、非线性、因果深度与通路结构差异;同时在真实世界场景中结合NHANES临床变量与外部生成的视网膜表征。结果表明,RetiSEM在合成数据上结构误差更低、因果准确性更高;真实数据中视网膜变量主要表现为下游生物标志物,具较小但可检测的间接效应。该策略为低资源医学人工智能中的结构化因果假设检验提供了可解释框架。代码与资源已公开于:https://github.com/Inamullah-Colab/ReitSEM。
原文摘要 · Abstract (English)
Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomplete or not jointly observed. We propose RetiSEM, a domain-constrained structural equation modelling (SEM) framework for causal graph recovery and mediation analysis under limited multimodal resources. This proposed work organises variables into biologically informed blocks, applies forbidden-edge constraints, and decomposes pathway-level effects into TE, NDE, and NIE components. We evaluate RetiSEM across ten synthetic benchmark scenarios that vary in dimensionality, nonlinearity, causal depth, and pathway structure, together with a fragmented real-world setting that combines NHANES clinical variables with externally derived retinal representations. This approach achieves lower structural error and higher causal accuracy than unconstrained baselines across the synthetic benchmarks. In the real-data analysis, retinal variables behave mainly as downstream biomarker-like indicators, with smaller but detectable indirect effects. These findings support our strategy as an interpretable framework for testing structured causal hypotheses in limited-resource biomedical AI. The code and resources for this work are publicly available at: https://github.com/Inamullah-Colab/ReitSEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。