通过对比健康人与病人,自动发现疾病亚型。
Automatic Discovery of Disease Subgroups by Contrasting with Healthy Controls

- 用深度模型对比病人与健康人,分离出疾病特异性特征。
- 在4个医学影像数据集上提升亚群识别质量。
- 适合做精准医疗或疾病分型的研究者使用。
在生物医学亚群发现中,研究者希望从患者群体中找出可解释且同质的亚组。本文假设健康个体(即对照组)与患者共享一些无关的变异因素,提出一种对比式亚群发现方法Deep UCSL。通过对比患者与对照组,该方法仅关注病理驱动的亚群,忽略与健康人共有的变异。框架采用深度特征提取器学习判别性表示空间,数学上推导出基于潜在聚类与患者/对照标签联合条件似然的新损失函数,通过期望最大化策略交替进行亚群推断和特征编码器更新。正则化项进一步促使表示捕捉疾病特异性变异,而非与对照组共享的变异。相比先前方法,该方法在MNIST示例及四个真实医学影像数据集上均显著提升亚群估计质量。代码与数据集见:https://github.com/rlouiset/deep_ucsl。
原文摘要 · Abstract (English)
In biomedical Subgroup Discovery, practitioners are interested in discovering interpretable and homogeneous subgroups within a group of patients. In this paper, assuming that healthy subjects (i.e., controls) share common but irrelevant factors of variation with the patients, we motivate and develop a Contrastive Subgroup Discovery method, entitled Deep UCSL. By contrasting patients with controls, Deep UCSL identifies subgroups driven solely by pathological factors, ignoring common variability shared with healthy subjects. Our framework employs a deep feature extractor to learn a discriminative representation space. Mathematically, we derive a novel loss based on the conditional joint likelihood of latent clusters and patient/control labels, optimized via an Expectation-Maximization strategy alternating between subgroup inference and feature encoder updates. A regularization term further encourages representations to capture disease-specific variability while ignoring variability shared with controls. Compared to previous related works, our approach quantitatively improves the quality of the estimated subgroups, as demonstrated on a MNIST example and four distinct real medical imaging datasets. Code and datasets are available at: https://github.com/rlouiset/deep_ucsl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。