将医学影像按器官拆分令牌,提升可解释性与临床应用灵活性。
OWT: A Foundational Organ-Wise Tokenization Framework for Medical Imaging
- 按器官划分图像令牌组,实现语义解耦
- 在CT/MRI上达成重建与分割优异性能
- 无需额外训练即可支持肿瘤定位等新功能
当前表示学习常依赖整体嵌入,混合多重语义信息,限制可解释性与泛化能力,尤其在医学影像中尤为关键。为此,我们提出器官级令牌化(OWT)框架,结合基于令牌组的重建(TGR)训练范式。不同于传统方法,OWT将图像显式分解为独立令牌组,每组对应一个器官或语义实体。该设计确保每个令牌组仅包含器官特异性信息,显著提升可解释性、泛化性与效率,并支持针对特定临床场景的细粒度控制。在CT和MRI数据集上的实验表明,OWT不仅在图像重建与分割等标准任务上表现优异,更实现了无需额外训练的器官级肿瘤识别、器官级检索与语义级生成等高影响力临床能力。这些发现凸显了OWT作为语义解耦表示学习基础框架的潜力,具备广泛可扩展性与新视角。
原文摘要 · Abstract (English)
Recent advances in representation learning often rely on holistic embeddings that entangle multiple semantic components, limiting interpretability and generalization. These issues are especially critical in medical imaging, where downstream tasks depend on anatomically interpretable features. To address these limitations, we propose an Organ-Wise Tokenization (OWT) framework with a Token Group-based Reconstruction (TGR) training paradigm. Unlike conventional approaches, OWT explicitly disentangles an image into separable token groups, each corresponding to a distinct organ or semantic entity. Our design ensures each token group encapsulates organ-specific information, boosting interpretability, generalization, and efficiency while enabling fine-grained control for targeted clinical applications. Experiments on CT and MRI datasets demonstrate OWT's power: it not only achieves strong performance on standard tasks like image reconstruction and segmentation, but also unlocks novel, high-impact clinical capabilities including organ-specific tumor identification, organ-level retrieval and semantic-level generation, without requiring any additional training. These findings underscore the potential of OWT as a foundational framework for semantically disentangled representation learning, offering broad scalability and a new perspective on how representations can be leveraged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。