让医学影像分割更懂人话,自动纠正表达差异
Skill-Evolving Grounded Reasoning for Free-Text Promptable 3D Medical Image Segmentation
- 用推理链将模糊描述与解剖结构对齐,确保语义准确
- 在语言扰动下性能方差降低81.94%,最差情况Dice提升18.60%
- 能自学习改进,适合临床场景中不规范的自由文本输入
自由文本提示的3D医学图像分割提供了直观且灵活的临床交互方式。然而,现有方法对语言变化极为敏感:微小措辞差异可能导致性能显著下降,即使临床意图相同。现有方法通过强化视觉-语言融合或扩大词汇量来提升鲁棒性,但缺乏将模糊表达与解剖学表征一致对齐的机制。本文提出技能演化式具身推理(SEER)框架,通过推理驱动设计显式弥合语言多样性与解剖精确性之间的差距。首先,构建了SEER-Trace数据集,将原始临床请求与图像引导的、带技能标签的推理轨迹配对,建立可复现基准。其次,SEER通过视觉-语言推理链构建证据对齐的目标表示,将临床意图与图像生成的解剖证据验证一致,从而在体素级解码前保障语义一致性。第三,提出SEER-Loop动态技能演化策略,将高收益推理路径提炼为可复用的技能模块,并逐步融入后续推理,实现结构化自我优化,提升对多样化语言表达的鲁棒性。大量实验表明,SEER优于当前最优基线。在语言扰动下,性能方差降低81.94%,最差情况下的Dice评分提升18.60%。
原文摘要 · Abstract (English)
Free-text promptable 3D medical image segmentation offers an intuitive and clinically flexible interaction paradigm. However, current methods are highly sensitive to linguistic variability: minor changes in phrasing can cause substantial performance degradation despite identical clinical intent. Existing approaches attempt to improve robustness through stronger vision-language fusion or larger vocabularies, yet they lack mechanisms to consistently align ambiguous free-form expressions with anatomically grounded representations. We propose Skill-Evolving grounded Reasoning (SEER), a novel framework for free-text promptable 3D medical image segmentation that explicitly bridges linguistic variability and anatomical precision through a reasoning-driven design. First, we curate the SEER-Trace dataset, which pairs raw clinical requests with image-grounded, skill-tagged reasoning traces, establishing a reproducible benchmark. Second, SEER constructs an evidence-aligned target representation via a vision-language reasoning chain that verifies clinical intent against image-derived anatomical evidence, thereby enforcing semantic consistency before voxel-level decoding. Third, we introduce SEER-Loop, a dynamic skill-evolving strategy that distills high-reward reasoning trajectories into reusable skill artifacts and progressively integrates them into subsequent inference, enabling structured self-refinement and improved robustness to diverse linguistic expressions. Extensive experiments demonstrate superior performance of SEER over state-of-the-art baselines. Under linguistic perturbations, SEER reduces performance variance by 81.94% and improves worst-case Dice by 18.60%. Project page: https://seer-medseg.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。