提出统一医学图像分词器,让模型更好理解与生成多模态医疗影像。
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
- 分两阶段训练:先用无标签影像学建立基础语义,再注入图文对细粒度信息
- 在3300万张跨模态医学图像和200万图文对上训练,性能超越30多个基准
- 适合需要融合诊断与生成能力的医学多模态研究者使用
自回归建模推动了多模态AI的重大进展,但其在医学影像领域的应用受限于缺乏能同时保留精细解剖结构与丰富临床语义的统一图像分词器。现有方法联合优化图像重建与文本语义目标,依赖大规模图像-文本配对数据,易受梯度干扰,而医学领域配对数据稀缺,大量未标注影像被浪费。本文识别该问题,提出一种基于视觉表征作为桥梁的两阶段训练框架。第一阶段利用大规模无配对医学影像进行视觉表征对齐,确保重建保真度并建立基础语义,缓解后续阶段的干扰;第二阶段使用图像-文本对注入细粒度语义。所提出的MedITok分词器在超过3300万张涵盖9种模态的医学图像及200万图像-文本对上训练,实现9种模态、4类任务共30+基准上的领先性能。它支持诊断与生成的自回归建模,可作为未来医学多模态模型中统一合成与理解的核心组件。
原文摘要 · Abstract (English)
Autoregressive modeling has driven major advances in multimodal AI, yet its application to medical imaging remains constrained by the absence of a unified image tokenizer that simultaneously preserves fine-grained anatomical structures and rich clinical semantics across heterogeneous modalities. Existing approaches jointly optimize image reconstruction and textual semantic objectives, relying on large-scale image-caption pairs and are prone to gradient interference. This is ill-suited for the medical domain where paired data are scarce and abundant unpaired images remain unexploited. This work identifies these issues in building unified medical image tokenizers, and introduces a principled two-stage training framework using visual representation as a bridge to address them. The propose visual representation alignment stage enables the utilization of large-scale unpaired medical images to ensure reconstruction fidelity and establish foundational semantics, alleviating the interference and better preparing for the second stage where fine-grained textual semantics are injected using image-text pairs. The resulting tokenizer, MedITok, is trained on over 33 million medical images spanning 9 modalities and 2 million image-text pairs. MedITok achieves state-of-the-art performance on 30+ benchmarks spanning 9 imaging modalities and 4 task families. It further enables autoregressive modeling for diagnostic and generative applications, serving as a scalable component for future multimodal models with unified synthesis and understanding capabilities in the medical domain. Project page: https://github.com/Masaaki-75/meditok
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。