arXiv:2510.10254cs.CV2025-10被引 3

无需医学数据训练,视频模型可直接零样本完成医学影像分割与运动预测。

Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?

  • 用自回归视频建模直接处理医学影像,不需微调。
  • 在122名患者1820个CT体积上实现竞争性分割、去噪与超分辨率效果。
  • 3D运动预测捕捉呼吸动态,时空一致性好,优于专用模型。

近期大型生成模型进展表明,适当规模的自回归模型可在跨领域任务中展现强大零样本泛化能力。受此启发,我们研究自回归视频建模是否可直接应用于医学影像任务,即使模型从未接触过医学数据。具体评估一个大视觉模型(LVM)在四类代表性任务上的零样本表现:器官分割、去噪、超分辨率和运动预测。令人惊讶的是,未经领域特定微调,该模型即可在CT扫描中勾画解剖结构,在分割、去噪和超分辨率任务上达到竞争性性能。尤为突出的是,在放疗运动预测任务中,模型能从4D CT扫描的先前相位直接预测未来3D CT相位,生成符合解剖结构且体现患者特异性呼吸动态的预测结果,具有真实的时间一致性。我们在122名患者的4D CT数据上进行评估,总计超过1,820个3D CT体积。尽管从未接触过医学数据,模型在所有任务中均表现优异,并在运动预测上超越基于形变场(DVF)及生成式基线模型,达到最先进的空间精度。这些发现揭示了医学视频建模中零样本能力的涌现,凸显通用视频模型作为统一学习者与推理者的潜力,为基于视频模型构建未来医学基础模型奠定基础。

原文摘要 · Abstract (English)

Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-shot generalization across domains. Motivated by this trend, we investigate whether autoregressive video modeling principles can be directly applied to medical imaging tasks, despite the model never being trained on medical data. Specifically, we evaluate a large vision model (LVM) in a zero-shot setting across four representative tasks: organ segmentation, denoising, super-resolution, and motion prediction. Remarkably, even without domain-specific fine-tuning, the LVM can delineate anatomical structures in CT scans and achieve competitive performance on segmentation, denoising, and super-resolution. Most notably, in radiotherapy motion prediction, the model forecasts future 3D CT phases directly from prior phases of a 4D CT scan, producing anatomically consistent predictions that capture patient-specific respiratory dynamics with realistic temporal coherence. We evaluate the LVM on 4D CT data from 122 patients, totaling over 1,820 3D CT volumes. Despite no prior exposure to medical data, the model achieves strong performance across all tasks and surpasses specialized DVF-based and generative baselines in motion prediction, achieving state-of-the-art spatial accuracy. These findings reveal the emergence of zero-shot capabilities in medical video modeling and highlight the potential of general-purpose video models to serve as unified learners and reasoners laying the groundwork for future medical foundation models built on video models.

视频生成医学影像零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。