提出因果感知的病理图像诊断框架,解决算法偏见与可解释性难题。
MeCaMIL: Causality-Aware Multiple Instance Learning for Fair and Interpretable Whole Slide Image Diagnosis
- 构建因果图模型,显式分离疾病信号与人口统计混淆因素。
- 在三个数据集上达到最高准确率,公平性提升超65%。
- 适合关注临床可解释性与公平性的数字病理研究者。
多实例学习(MIL)已成为计算病理学中全切片图像(WSI)分析的主流范式,通过局部特征聚合实现优异诊断性能。然而现有方法存在两大局限:一是依赖缺乏因果可解释性的注意力机制;二是未整合患者年龄、性别、种族等人口统计信息,导致跨群体公平性问题。本文提出MeCaMIL,一种基于结构化因果图的因果感知MIL框架,利用do-calculus和碰撞器结构,显式建模人口统计混淆因子,分离疾病相关信号与虚假关联。在三个基准测试中表现卓越:CAMELYON16(ACC/AUC/F1: 0.939/0.983/0.946)、TCGA-Lung(0.935/0.979/0.931)、TCGA-Multi(0.977/0.993/0.970,五种癌症)。关键成果:人口统计差异方差平均降低65%以上,对弱势群体改善显著;生存预测平均C-index达0.653,优于最优基线0.017。消融实验验证因果图结构必要性——替代设计导致准确率下降0.048,公平性恶化4.2倍。结果证明该框架为数字病理中公平、可解释、可临床落地的AI提供新范式。
原文摘要 · Abstract (English)
Multiple instance learning (MIL) has emerged as the dominant paradigm for whole slide image (WSI) analysis in computational pathology, achieving strong diagnostic performance through patch-level feature aggregation. However, existing MIL methods face critical limitations: (1) they rely on attention mechanisms that lack causal interpretability, and (2) they fail to integrate patient demographics (age, gender, race), leading to fairness concerns across diverse populations. These shortcomings hinder clinical translation, where algorithmic bias can exacerbate health disparities. We introduce \textbf{MeCaMIL}, a causality-aware MIL framework that explicitly models demographic confounders through structured causal graphs. Unlike prior approaches treating demographics as auxiliary features, MeCaMIL employs principled causal inference -- leveraging do-calculus and collider structures -- to disentangle disease-relevant signals from spurious demographic correlations. Extensive evaluation on three benchmarks demonstrates state-of-the-art performance across CAMELYON16 (ACC/AUC/F1: 0.939/0.983/0.946), TCGA-Lung (0.935/0.979/0.931), and TCGA-Multi (0.977/0.993/0.970, five cancer types). Critically, MeCaMIL achieves superior fairness -- demographic disparity variance drops by over 65% relative reduction on average across attributes, with notable improvements for underserved populations. The framework generalizes to survival prediction (mean C-index: 0.653, +0.017 over best baseline across five cancer types). Ablation studies confirm causal graph structure is essential -- alternative designs yield 0.048 lower accuracy and 4.2x times worse fairness. These results establish MeCaMIL as a principled framework for fair, interpretable, and clinically actionable AI in digital pathology. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。