用眼球追踪替代鼠标点击,实现更高效的眼动交互式医学图像分割。
Zero-Shot Gaze-based Volumetric Medical Image Segmentation
- 以眼球轨迹作为交互提示,替代传统鼠标点击或框选。
- 相比框选,眼动提示效率更高但精度略低,实时性提升显著。
- 适合需要快速标注的临床场景,如手术规划与疾病监测。
在三维医学图像中精确分割解剖结构对疾病监测和癌症治疗规划至关重要。当前交互式分割模型(如SAM-2及其医学变体MedSAM-2)依赖人工提供的提示,如边界框和鼠标点击。本研究首次将眼动追踪引入3D医学图像分割,提出以眼球运动作为新型交互模态。我们在合成与真实眼动数据上评估了眼动提示在SAM-2和MedSAM-2上的表现。结果表明,相较于边界框,眼动提示在交互效率上更具优势,虽分割精度略有下降,但显著提升了操作速度。研究证实眼动可作为交互式3D医学图像分割的互补输入方式。
原文摘要 · Abstract (English)
Accurate segmentation of anatomical structures in volumetric medical images is crucial for clinical applications, including disease monitoring and cancer treatment planning. Contemporary interactive segmentation models, such as Segment Anything Model 2 (SAM-2) and its medical variant (MedSAM-2), rely on manually provided prompts like bounding boxes and mouse clicks. In this study, we introduce eye gaze as a novel informational modality for interactive segmentation, marking the application of eye-tracking for 3D medical image segmentation. We evaluate the performance of using gaze-based prompts with SAM-2 and MedSAM-2 using both synthetic and real gaze data. Compared to bounding boxes, gaze-based prompts offer a time-efficient interaction approach with slightly lower segmentation quality. Our findings highlight the potential of using gaze as a complementary input modality for interactive 3D medical image segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。