统一病理与影像的医学多任务基础模型,支持跨模态、多粒度预测。
CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

- 通过注意力机制建模多维上下文,融合病理与影像数据
- 在5个断层扫描、4个全组织、3个二维数据集上表现领先
- 仅微调2.5%参数即达全量微调效果,适合资源受限研究
医学基础模型在标注数据有限时可提升AI模型泛化能力,但通常局限于单一专科(如病理科或放射科)及单一输出形式(如分类或分割)。本文提出CoM$^3$eT(联合表征多维度多任务医学Transformer),首次统一病理科与放射科数据,支持稀疏与密集输出,并处理二维及以上输入,通过注意力机制建模多维上下文。在涵盖五种断层扫描、四种全组织和三种二维数据集的开放竞赛中表现超越现有医学基础模型,覆盖分类、分割及报告生成任务。模型适应多种临床场景时,仅需微调少于2.5%参数即可达到全量微调性能,无需高性能GPU集群。在跨医院联邦学习中,该方法在互联网连接与消费级硬件条件下实现与集中式训练相当的效果。
原文摘要 · Abstract (English)
Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or radiology, and to either sparse or dense outputs, such as classification or segmentation. Here, we present CoM$^3$eT (Co-representation Multidimensional Multitask Medical Transformer), a medical vision foundation model that unifies pathology and radiology, sparse and dense predictions, and two- and higher-dimensional inputs by modeling multidimensional context with attention. CoM$^3$eT outperformed other medical foundation models in an open competition spanning five tomographic, four whole-specimen, and three two-dimensional datasets, covering sparse and dense prediction tasks as well as report generation. When adapted across diverse clinical applications, training fewer than 2.5% of parameters achieved performance comparable to full fine-tuning, enabling research without access to high-performance GPU clusters. Applied to federated learning across hospitals, this approach achieved performance comparable to pooled-data training over internet connections and with consumer-grade hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。