用生成模型衡量患者与训练数据的相似度,提升外部验证的解释力。
Rethinking external validation for the target population: Capturing patient-level similarity with a generative model
- 通过自编码器量化外部患者与原始数据的相似性,替代传统线性方法。
- 发现模型在不同相似度子组表现差异大,常规验证可能掩盖关键问题。
- 适合关注模型可迁移性及临床应用安全性的研究人员与医生。
背景:外部验证对评估预测模型的可迁移性至关重要,但常因外部与开发人群差异而难以解读。本研究提出一种框架,区分模型缺陷与病例组合效应。方法:利用自编码器估计每位外部患者与开发数据的相似性,按相似度分层评估性能,无需共享原始开发数据。通过模拟数据验证自编码器相似性度量的有效性,并以荷兰心脏注册(NHR)数据为例,预测经导管主动脉瓣植入术后的死亡率。结果:框架显示,不同相似度子组间模型表现存在显著差异,常规外部验证未能揭示这些差异,却可能影响结论。某些情况下,常规验证显示整体性能差,但经调整后,部分子组表现与内部验证一致;反之,看似可接受的整体性能可能掩盖特定子组的临床相关缺陷。结论:该框架通过关联模型性能与人群分布对齐程度,增强外部验证的可解释性,为判断模型是否可迁移及适用于哪些患者提供更严谨依据。
原文摘要 · Abstract (English)
Background: External validation is essential for assessing the transportability of predictive models. However, its interpretation is often confounded by differences between external and development populations. This study introduces a framework to distinguish model deficiencies from case-mix effects. Method: We propose a framework that quantifies each external patient's similarity to the development data and measures performance in subgroups with varying levels of alignment to the development distribution. We use generative models, specifically autoencoders, to estimate similarity, offering a more flexible alternative to traditional linear approaches and enabling validation without sharing the original development data. The utility of autoencoder-based similarity measure is demonstrated using synthetic data, and the framework's application is illustrated using data from the Netherlands Heart Registration (NHR) to predict mortality after transcatheter aortic valve implantation. Results: Our framework revealed substantial variation in model performance across similarity-defined subgroups, differences that remain hidden under conventional external validation yet can meaningfully alter conclusions. In several settings, conventional external validation suggested poor overall performance. However, after accounting for differences in patient characteristics, for some sub-groups, the model performance was consistent with internal validation results. Conversely, apparently acceptable overall performance could mask clinically relevant performance deficits in specific subgroups. Conclusion: The proposed framework enhances the interpretability of external validation by linking model performance to population alignment with the development data. This provides a more principled basis for deciding whether a model is transportable and to which patients it can be safely applied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。