arXiv:2604.05620cs.CVcs.AI2026-04被引 1

用图推理让AI理解肺部影像报告,精准定位病灶。

Semantic-Topological Graph Reasoning for Language-Guided Pulmonary Screening

  • 将病变视为节点,通过空间与语义关系建图,解决重叠模糊问题。
  • 在LIDC-IDRI数据集上达81.5%的分割准确率,优于主流工具超5%。
  • 仅更新不足1%参数,适合临床部署且跨验证稳定性强。

基于自由文本临床指令的医学图像分割是辅助诊断的关键前沿。然而,现有多模态与基础模型难以处理报告中的语义歧义,在低对比度扫描中无法有效区分复杂解剖重叠。此外,在有限医疗数据集上全量微调这些大型架构极易导致严重过拟合。为此,我们提出一种新型语义-拓扑图推理(STGR)框架,用于语言引导的肺部筛查。该方法巧妙融合大语言模型(LLaMA-3-V)的推理能力与视觉基础模型(MedSAM)的零样本分割能力。具体地,引入文本到视觉意图蒸馏(TVID)模块提取精确诊断指引;为解决解剖歧义,将掩码选择建模为动态图推理问题,候选病灶作为节点,边表示空间与语义相似性。为保障部署可行性,提出选择性非对称微调(SAFT)策略,仅更新少于1%的参数。在LIDC-IDRI和LNDb数据集上的五折交叉验证表明,该框架达到新最优性能。尤其在LIDC-IDRI上实现81.5%的骰子相似系数(DSC),较领先工具如LISA高出5%以上。关键的是,SAFT策略起强大正则化作用,跨折稳定性优异(DSC方差仅0.6%),为鲁棒、上下文感知的临床应用铺平道路。

原文摘要 · Abstract (English)

Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle with the semantic ambiguity of clinical reports and fail to disambiguate complex anatomical overlaps in low-contrast scans. Furthermore, fully fine-tuning these massive architectures on limited medical datasets invariably leads to severe overfitting. To address these challenges, we propose a novel Semantic-Topological Graph Reasoning (STGR) framework for language-guided pulmonary screening. Our approach elegantly synergizes the reasoning capabilities of large language models (LLaMA-3-V) with the zero-shot delineation of vision foundation models (MedSAM). Specifically, we introduce a Text-to-Vision Intent Distillation (TVID) module to extract precise diagnostic guidance. To resolve anatomical ambiguity, we formulate mask selection as a dynamic graph reasoning problem, where candidate lesions are modeled as nodes and edges capture spatial and semantic affinities. To ensure deployment feasibility, we introduce a Selective Asymmetric Fine-Tuning (SAFT) strategy that updates less than 1% of the parameters. Rigorous 5-fold cross-validation on the LIDC-IDRI and LNDb datasets demonstrates that our framework establishes a new state-of-the-art. Notably, it achieves an 81.5% Dice Similarity Coefficient (DSC) on LIDC-IDRI, outperforming leading LLM-based tools like LISA by over 5%. Crucially, our SAFT strategy acts as a powerful regularizer, yielding exceptional cross-fold stability (0.6% DSC variance) and paving the way for robust, context-aware clinical deployment.

肺部筛查图推理多模态小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。