arXiv:2602.14512cs.CV2026-02被引 2

MedVAR用分层预测生成医学影像,速度快且可扩展。

MedVAR: Towards Scalable and Efficient Medical Image Generation via Next-scale Autoregressive Prediction

  • 采用逐尺度自回归生成,从粗到细构建多尺度图像
  • 在约44万张CT/MRI数据上训练,支持六部位图像生成
  • 生成质量与扩展性俱佳,适合医疗数据增强和隐私共享

医学图像生成在低资源临床任务的数据增强和隐私保护数据共享中至关重要。然而,构建可扩展的医学图像生成基础模型需兼顾架构效率、充足多器官数据及合理评估,现有方法尚未解决这些问题。为此,我们提出MedVAR,首个基于自回归的基座模型,采用下一尺度预测范式,实现快速且可扩展的医学图像合成。MedVAR以粗到细的方式生成图像,并生成适用于下游任务的结构化多尺度表征。为支持层级生成,我们整理了一个包含约44万张涵盖六个解剖区域的CT和MRI图像的统一数据集。在保真度、多样性与可扩展性方面的综合实验表明,MedVAR达到当前最优生成性能,为未来医学生成基座模型提供了有前景的架构方向。

原文摘要 · Abstract (English)

Medical image generation is pivotal in applications like data augmentation for low-resource clinical tasks and privacy-preserving data sharing. However, developing a scalable generative backbone for medical imaging requires architectural efficiency, sufficient multi-organ data, and principled evaluation, yet current approaches leave these aspects unresolved. Therefore, we introduce MedVAR, the first autoregressive-based foundation model that adopts the next-scale prediction paradigm to enable fast and scale-up-friendly medical image synthesis. MedVAR generates images in a coarse-to-fine manner and produces structured multi-scale representations suitable for downstream use. To support hierarchical generation, we curate a harmonized dataset of around 440,000 CT and MRI images spanning six anatomical regions. Comprehensive experiments across fidelity, diversity, and scalability show that MedVAR achieves state-of-the-art generative performance and offers a promising architectural direction for future medical generative foundation models.

医学影像自回归生成多尺度数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。