用扩散模型生成脑影像数据,提升精神分裂症分类准确率
MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model
- 基于潜在扩散模型生成增强脑影像数据,缓解过拟合
- 新模型在精神分裂症分类上比旧版准确率显著提升
- 适合需要高泛化能力的多模态医学影像研究者
多模态医学影像融合结构与功能脑影像数据,为深度学习预测提供互补信息,提升诊断效果。本研究提出新一代神经影像预测框架 MultiViT2,结合预训练表示学习基模型与视觉变换器骨干网络进行预测输出。此外,开发了基于潜在扩散模型的数据增强模块,通过生成增强的脑影像样本,降低过拟合,提升模型泛化能力。实验表明,MultiViT2 在精神分裂症分类任务中显著优于第一代模型,具备强可扩展性与可迁移性。
原文摘要 · Abstract (English)
Multimodal medical imaging integrates diverse data types, such as structural and functional neuroimaging, to provide complementary insights that enhance deep learning predictions and improve outcomes. This study focuses on a neuroimaging prediction framework based on both structural and functional neuroimaging data. We propose a next-generation prediction model, \textbf{MultiViT2}, which combines a pretrained representative learning base model with a vision transformer backbone for prediction output. Additionally, we developed a data augmentation module based on the latent diffusion model that enriches input data by generating augmented neuroimaging samples, thereby enhancing predictive performance through reduced overfitting and improved generalizability. We show that MultiViT2 significantly outperforms the first-generation model in schizophrenia classification accuracy and demonstrates strong scalability and portability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。