用合成数据模拟眼扫描与视力变化关系,可控关联水平助力模型评估。
A Framework for Evaluating Predictive Models Using Synthetic Image Covariates and Longitudinal Data
- 在潜在空间控制多模态数据关联,生成带时间序列的医疗图像
- 110万张合成眼扫图与5组不同关联度(2%-100%)的视力数据配对
- 可检测微弱信号,适合隐私敏感的医疗模型验证与基准测试
我们提出一种新框架,通过在潜在空间中控制关联性,合成复杂协变量(如眼部扫描)与纵向观测(如随时间变化的视力)的数据对,解决医疗研究中的隐私问题。该方法结合变分自编码器与扩散模型,基于109,309张二维OCT扫描切片训练图像生成模型;纵向数据则通过非线性混合效应(NLME)模型从低维随机效应空间生成。共生成110万张配对的OCT切片,对应五种不同关联水平(100%、50%、10%、5.26%和2%的个体间变异)。为评估框架有效性,我们使用另一NLME模型拟合合成纵向数据,计算随机效应的经验贝叶斯估计,并训练一个ResNet模型从合成OCT图像预测这些估计值。随后将预测结果嵌入原NLME模型实现个性化预测。结果显示,随着图像与观测间关联度下降,预测精度按预期降低;但在除2%外的所有情况下,预测性能均达到理论最优值的50%以内,证明框架能有效捕捉微弱信号。该框架可生成具有可控关联性的合成数据,为医疗研究提供宝贵工具。
原文摘要 · Abstract (English)
We present a novel framework for synthesizing patient data with complex covariates (e.g., eye scans) paired with longitudinal observations (e.g., visual acuity over time), addressing privacy concerns in healthcare research. Our approach introduces controlled association in latent spaces generating each data modality, enabling the creation of complex covariate-longitudinal observation pairs. This framework facilitates the development of predictive models and provides openly available benchmarking datasets for healthcare research. We demonstrate our framework using optical coherence tomography (OCT) scans, though it is applicable across domains. Using 109,309 2D OCT scan slices, we trained an image generative model combining a variational autoencoder and a diffusion model. Longitudinal observations were simulated using a nonlinear mixed effect (NLME) model from a low-dimensional space of random effects. We generated 1.1M OCT scan slices paired with five sets of longitudinal observations at controlled association levels (100%, 50%, 10%, 5.26%, and 2% of between-subject variability). To assess the framework, we modeled synthetic longitudinal observations with another NLME model, computed empirical Bayes estimates of random effects, and trained a ResNet to predict these estimates from synthetic OCT scans. We then incorporated ResNet predictions into the NLME model for patient-individualized predictions. Prediction accuracy on withheld data declined as intended with reduced association between images and longitudinal measurements. Notably, in all but the 2% case, we achieved within 50% of the theoretical best possible prediction on withheld data, demonstrating our ability to detect even weak signals. This confirms the effectiveness of our framework in generating synthetic data with controlled levels of association, providing a valuable tool for healthcare research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。