用语音指令逐步修正3D医学图像分割结果,提升临床可用性。
Refining 3D Medical Segmentation with Verbal Instruction
- 将3D解剖结构表示为向量集,通过文本指令迭代优化形状
- 在合成错误数据上实现显著精度提升,优于现有基线方法
- 适合需要医生参与的医疗影像精细分割场景
精准的3D解剖分割对临床诊断和手术规划至关重要。但自动化模型常因训练数据有限且不均衡、标注质量差以及训练与部署分布差异,产生次优形状预测。自然解决方案是基于放射科医生的语音指令,迭代修正预测形状。然而,这一方法受限于配对数据稀缺——即错误形状与其对应修正指令的明确关联。为此,我们提出CoWTalk基准,包含可控制合成的3D动脉解剖错误及其修复指令。基于该基准,我们进一步设计一种迭代精修模型,将3D形状表示为向量集合,并与文本指令交互以逐步更新目标形状。实验表明,该方法在噪声输入下显著优于原始结果和竞争基线,验证了语言驱动的临床医生协同修正3D医学形态建模的可行性。
原文摘要 · Abstract (English)
Accurate 3D anatomical segmentation is essential for clinical diagnosis and surgical planning. However, automated models frequently generate suboptimal shape predictions due to factors such as limited and imbalanced training data, inadequate labeling quality, and distribution shifts between training and deployment settings. A natural solution is to iteratively refine the predicted shape based on the radiologists' verbal instructions. However, this is hindered by the scarcity of paired data that explicitly links erroneous shapes to corresponding corrective instructions. As an initial step toward addressing this limitation, we introduce CoWTalk, a benchmark comprising 3D arterial anatomies with controllable synthesized anatomical errors and their corresponding repairing instructions. Building on this benchmark, we further propose an iterative refinement model that represents 3D shapes as vector sets and interacts with textual instructions to progressively update the target shape. Experimental results demonstrate that our method achieves significant improvements over corrupted inputs and competitive baselines, highlighting the feasibility of language-driven clinician-in-the-loop refinement for 3D medical shapes modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。