arXiv:2608.05389cs.CV2026-08

用文本指令精准修正脑瘤分割结果,提升放疗规划准确性

Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

论文配图:Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model
图 1 · 摘自论文原文
  • 用3D视觉语言模型注入文本指令,动态调整分割边界
  • 正确指令使肿瘤分割精度提升至0.796(原0.774)
  • 支持医生通过自然语言直接修改结果,适合临床辅助

准确划分胶质瘤亚区对放疗计划与长期监测至关重要,但人工勾画耗时。现有模型如nnU-Net泛化能力不足且缺乏医生引导的文本修正机制。本文探索将三维视觉语言基础模型用于文本引导的脑肿瘤分割精修。提出轻量级VoxTell框架:预训练VoxTell生成初始分割掩膜;基于分割误差构建包含目标、动作、位置、影像证据、编辑范围和保留约束的精确提示。冻结的Qwen/VoxTell提示嵌入通过可训练投影注入多尺度解码器条件模块,其余参数保持冻结。训练、验证与测试分别使用901、100和250例BraTS-GLI数据。跨数据集迁移在100例脑膜瘤、转移瘤、儿童肿瘤及UPENN-GBM病例中评估。结果表明,在内部测试集上,使用增强后对比T1加权图像输入,正确指令使子区Dice相似系数(DSC)从0.774±0.158提升至0.796±0.137;显著优于空白提示(0.762±0.155;霍尔姆校正p<0.001,d_z=0.71)与矛盾提示(0.770±0.163;p<0.001,d_z=0.48)。跨数据集测试中,正确指令使DSC从0.527±0.287升至0.550±0.278,并优于矛盾指令(0.504±0.275;p<0.001,d_z=0.43)。结论:3D视觉语言基础模型可实现指令引导的胶质瘤亚区分割精修。对正确、空白及矛盾提示的敏感性表明其为文本依赖性编辑,而非非特定后处理,支持其作为医生在环工具的进一步评估。

原文摘要 · Abstract (English)

Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such as nnU-Net may generalize imperfectly and lack clinician-directed text correction. Purpose: We investigated adapting a three-dimensional (3D) vision-language foundation model for text-guided brain tumor segmentation refinement. Methods: We developed a lightweight VoxTell-based framework. Pretrained VoxTell generated initial masks. Oracle prompts derived from segmentation errors encoded target, action, location, imaging evidence, edit size, and preservation constraints. Frozen Qwen/VoxTell prompt embeddings were injected through trainable projections into its multiscale decoder conditioning; other weights remained frozen. Training, validation, and testing used 901, 100, and 250 BraTS-GLI cases. Cross-dataset transfer was evaluated on 100 meningioma, metastasis, pediatric tumor, and UPENN-GBM cases. Results: On the internal test set using post-contrast T1-weighted input, correct instructions improved subregion Dice similarity coefficient (DSC; enhancing tumor, edema, and necrotic/non-enhancing core) from $0.774\pm0.158$ to $0.796\pm0.137$. They outperformed blank prompts ($0.762\pm0.155$; Holm-adjusted $p<0.001$, $d_z=0.71$) and contradictory prompts ($0.770\pm0.163$; $p<0.001$, $d_z=0.48$). In cross-dataset testing, correct instructions improved DSC from $0.527\pm0.287$ to $0.550\pm0.278$ and outperformed contradictory instructions ($0.504\pm0.275$; $p<0.001$, $d_z=0.43$). Conclusion: A 3D vision-language foundation model can perform instruction-guided refinement of glioma subregion segmentations. Sensitivity to correct, blank, and contradictory prompts suggests text-dependent contour editing rather than nonspecific post-processing, supporting further evaluation as a clinician-in-the-loop tool.

医学影像文本引导分割精修视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。