arXiv:2602.16110cs.CVcs.AI2026-02被引 7

统一切片与三维理解,提升CT影像诊断精度

OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis

  • 融合切片与体积信息,通过位置嵌入和混合专家模型实现空间一致性
  • 在多种临床任务中显著超越现有方法,同时兼顾细节与宏观结构分析
  • 适合医学影像研究者及临床辅助诊断系统开发者使用

计算机断层扫描(CT)是应用最广泛、信息密度最高的影像模态之一,涵盖心脏、肺、肝、结肠等关键器官。临床解读依赖于切片级局部特征(如亚厘米结节、病灶边界)和体积级空间表征(如肿瘤浸润、器官间解剖关系)。然而,现有大型视觉语言模型(LVLM)在切片与体积理解上仍呈碎片化:切片驱动的模型泛化能力强但缺乏跨切片空间一致性;体积驱动的模型能捕捉体积语义但粒度粗糙且不兼容切片输入。缺乏统一建模范式已成为医疗LVLM临床落地的主要瓶颈。本文提出OmniCT,一种面向CT场景的统一切片-体积LVLM,包含三项贡献:(i) 空间一致性增强(SCE):通过体积切片组合与三轴位置嵌入引入体积一致性,并采用MoE混合投影实现高效切片-体积适配;(ii) 器官级语义增强(OSE):显式对齐解剖区域,强化病灶与器官级语义;(iii) MedEval-CT:目前最大的切片-体积联合数据集与混合基准,集成全面评估指标。OmniCT在多种临床任务中持续优于现有方法,显著提升微尺度细节敏感性与宏观空间推理能力。更重要的是,它建立了跨模态医学影像理解的新范式。项目代码已开源:https://github.com/ZJU4HealthCare/OmniCT。

原文摘要 · Abstract (English)

Computed Tomography (CT) is one of the most widely used and diagnostically information-dense imaging modalities, covering critical organs such as the heart, lungs, liver, and colon. Clinical interpretation relies on both slice-driven local features (e.g., sub-centimeter nodules, lesion boundaries) and volume-driven spatial representations (e.g., tumor infiltration, inter-organ anatomical relations). However, existing Large Vision-Language Models (LVLMs) remain fragmented in CT slice versus volumetric understanding: slice-driven LVLMs show strong generalization but lack cross-slice spatial consistency, while volume-driven LVLMs explicitly capture volumetric semantics but suffer from coarse granularity and poor compatibility with slice inputs. The absence of a unified modeling paradigm constitutes a major bottleneck for the clinical translation of medical LVLMs. We present OmniCT, a powerful unified slice-volume LVLM for CT scenarios, which makes three contributions: (i) Spatial Consistency Enhancement (SCE): volumetric slice composition combined with tri-axial positional embedding that introduces volumetric consistency, and an MoE hybrid projection enables efficient slice-volume adaptation; (ii) Organ-level Semantic Enhancement (OSE): segmentation and ROI localization explicitly align anatomical regions, emphasizing lesion- and organ-level semantics; (iii) MedEval-CT: the largest slice-volume CT dataset and hybrid benchmark integrates comprehensive metrics for unified evaluation. OmniCT consistently outperforms existing methods with a substantial margin across diverse clinical tasks and satisfies both micro-level detail sensitivity and macro-level spatial reasoning. More importantly, it establishes a new paradigm for cross-modal medical imaging understanding. Our project is available at https://github.com/ZJU4HealthCare/OmniCT.

医学影像视觉语言模型三维理解CT分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。