arXiv:2603.13054cs.CV2026-03被引 2

用视觉语言模型识别管状结构中的拓扑异常,提升医疗影像分析准确性。

Topo-R1: Detecting Topological Anomalies via Vision-Language Models

  • 构建多领域合成数据集,标注拓扑异常与贝蒂数变化。
  • 提出复合奖励机制,联合优化定位、分类与骨架结构保真度。
  • 在真实分割输出上表现超越主流模型,适合医学图像分析场景。

管状结构(如血管、神经纤维、道路网络)的拓扑特性对功能分析至关重要,其连通性和环状结构起决定作用。尽管视觉语言模型(VLMs)具备推理和定位能力,但我们在局部化和分类四类典型拓扑异常(断裂/虚假连接、缺失/多余分支)的分割掩码上系统评估了主流闭源与开源VLMs,发现其表现接近随机,表明当前通用VLM缺乏拓扑感知能力。由于缺乏带局部异常标注的分割掩码资源,我们构建了自动化多领域数据采集管道,通过可验证贝蒂数(Betti-number)标注生成不同难度层级的拓扑扰动数据,创建首个大规模训练集及分布内(ID)、分布外(OOD)测试集组成的系统性基准。基于此基准,我们提出Topo-R1:以拓扑感知复合奖励为核心,联合评分定位、分类与骨架结构保真度。通过监督微调初始化符合模式的输出,并采用组相对策略优化(GRPO)针对该奖励优化策略,引导预测趋向拓扑合理结构而非仅像素重叠。大量实验表明,Topo-R1显著优于通用VLMs,在ID、OOD及真实分割输出协议下达到或超过监督基线,为基于VLM的结构化视觉数据拓扑理解奠定坚实基础。

原文摘要 · Abstract (English)

Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern downstream functional analysis. Vision-Language Models (VLMs) are promising candidates for understanding such structures, given their reasoning and grounding capabilities. To probe their topological perception, we systematically evaluate leading closed- and open-source VLMs on localizing and classifying four canonical topological anomalies (broken/spurious connections, missing/extra branches) in tubular-network segmentation masks. They perform nearly at random, indicating that topology-aware perception is largely absent from current general-purpose VLMs. As no existing resource pairs segmentation masks with localized anomaly annotations, we build an automated, multi-domain data-curation pipeline that synthesizes diverse topological perturbations with verifiable Betti-number annotations across graduated difficulty levels, yielding the first systematic benchmark with a large-scale training set and held-out in-distribution (ID) and out-of-distribution (OOD) test suites. Building on this benchmark, we introduce Topo-R1, centered on a topology-aware composite reward that jointly scores localization, classification, and skeleton-level structural fidelity. Supervised fine-tuning cold-starts schema-compliant outputs, and Group Relative Policy Optimization (GRPO) then optimizes the policy against this reward, steering predictions toward topologically meaningful structure rather than superficial pixel overlap. Extensive experiments show that Topo-R1 substantially outperforms general-purpose VLMs and matches or exceeds supervised baselines across ID, OOD, and real-segmentation-output protocols, establishing a strong foundation for VLM-based topological understanding of structured visual data.

拓扑检测视觉语言模型医学影像数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。