arXiv:2603.09217cs.CV2026-03

用大模型理解血管拓扑,解决分割中断连和伪连接问题。

TubeMLLM: A Foundation Model for Topology Knowledge Exploration in Vessel-like Anatomy

  • 通过自然语言提示注入拓扑先验,结合视觉特征实现统一建模。
  • 在跨模态任务中零样本迁移效果显著,血管分割Dice达67.50%。
  • 适合医学图像分析、血管结构建模及跨模态研究者使用。

由于血管样解剖结构的复杂拓扑特性及其对数据分布偏移的敏感性,传统任务特定模型常出现拓扑不一致问题,如人为断连和虚假合并。为此,我们提出TubeMLLM,一种统一的基础模型,融合结构化理解与可控生成能力以应对医学血管结构建模挑战。通过显式自然语言提示引入拓扑先验,并在共享注意力架构中对齐视觉表征,显著提升拓扑感知能力。我们构建了首个多模态基准TubeMData,涵盖全面的拓扑中心任务,并设计自适应损失加权策略以强化拓扑关键区域训练。在十五个多样化数据集上的实验表明,该模型在分布外场景下表现领先:在彩色眼底照片上,$β_{0}$ 数量误差从37.42降至8.58;在未见过的X射线血管造影中实现67.50%的Dice分数,$β_{0}$误差降至1.21。模型对模糊、噪声、低分辨率等退化也表现出强鲁棒性。在拓扑质量评估任务中,准确率达97.38%,远超标准视觉-语言基线。

原文摘要 · Abstract (English)

Modeling medical vessel-like anatomy is challenging due to its intricate topology and sensitivity to dataset shifts. Consequently, task-specific models often suffer from topological inconsistencies, including artificial disconnections and spurious merges. Motivated by the promise of multimodal large language models (MLLMs) for zero-shot generalization, we propose TubeMLLM, a unified foundation model that couples structured understanding with controllable generation for medical vessel-like anatomy. By integrating topological priors through explicit natural language prompting and aligning them with visual representations in a shared-attention architecture, TubeMLLM significantly enhances topology-aware perception. Furthermore, we construct TubeMData, a pionner multimodal benchmark comprising comprehensive topology-centric tasks, and introduce an adaptive loss weighting strategy to emphasize topology-critical regions during training. Extensive experiments on fifteen diverse datasets demonstrate our superiority. Quantitatively, TubeMLLM achieves state-of-the-art out-of-distribution performance, substantially reducing global topological discrepancies on color fundus photography (decreasing the $β_{0}$ number error from 37.42 to 8.58 compared to baselines). Notably, TubeMLLM exhibits exceptional zero-shot cross-modality transferring ability on unseen X-ray angiography, achieving a Dice score of 67.50% while significantly reducing the $β_{0}$ error to 1.21. TubeMLLM also maintains robustness against degradations such as blur, noise, and low resolution. Furthermore, in topology-aware understanding tasks, the model achieves 97.38% accuracy in evaluating mask topological quality, significantly outperforming standard vision-language baselines.

医学图像拓扑建模大模型血管分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。