用可控合成数据评估临床模型跨机构泛化能力
Bridging the Generalisation Gap: Synthetic Data Generation for Multi-Site Clinical Model Validation
- 构建可调控的结构化合成数据框架,控制站点差异和子群效应
- 揭示模型复杂度与机构差异的交互影响,暴露泛化缺陷
- 适合临床模型开发者用于公平性审计和鲁棒性测试
临床机器学习模型在不同医疗场景中的泛化能力仍面临挑战,主要源于患者人口统计、疾病流行率及机构实践的差异。现有评估方法依赖真实数据,存在获取难、存在混杂偏差且难以系统实验的问题。尽管生成模型追求统计真实性,但缺乏透明度和对分布偏移驱动因素的显式控制。本文提出一种新型结构化合成数据框架,旨在实现模型鲁棒性、公平性和泛化能力的可控基准测试。不同于仅模仿观测数据的方法,本框架可显式控制数据生成过程,包括站点特异性流行率变化、层级子群效应及结构化特征交互。通过受控实验,验证了该框架能隔离站点差异的影响,支持公平性审计,并揭示泛化失败,尤其凸显模型复杂度与站点特异性效应的相互作用。本研究提供了一个可复现、可解释、可配置的工具,助力临床机器学习的可靠部署。
原文摘要 · Abstract (English)
Ensuring the generalisability of clinical machine learning (ML) models across diverse healthcare settings remains a significant challenge due to variability in patient demographics, disease prevalence, and institutional practices. Existing model evaluation approaches often rely on real-world datasets, which are limited in availability, embed confounding biases, and lack the flexibility needed for systematic experimentation. Furthermore, while generative models aim for statistical realism, they often lack transparency and explicit control over factors driving distributional shifts. In this work, we propose a novel structured synthetic data framework designed for the controlled benchmarking of model robustness, fairness, and generalisability. Unlike approaches focused solely on mimicking observed data, our framework provides explicit control over the data generating process, including site-specific prevalence variations, hierarchical subgroup effects, and structured feature interactions. This enables targeted investigation into how models respond to specific distributional shifts and potential biases. Through controlled experiments, we demonstrate the framework's ability to isolate the impact of site variations, support fairness-aware audits, and reveal generalisation failures, particularly highlighting how model complexity interacts with site-specific effects. This work contributes a reproducible, interpretable, and configurable tool designed to advance the reliable deployment of ML in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。