预测手术工具与组织的交互区域,提升自动化手术安全性和精准度。
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
- 融合视觉、语言和时间动态,生成工具-动作特异的组织可操作性热图。
- 在103例胆囊切除术中实现20.6像素平均表面距离,显著优于基线模型。
- 适用于需要高精度安全控制的智能手术机器人研发人员。
手术自动化正朝着类外科医生的灵巧控制快速演进,主要得益于从示范学习和视觉-语言-动作模型的进步。尽管这些方法在桌面实验中表现良好,但向临床部署转化仍面临挑战:现有方法对器械在组织表面的交互位置预测能力有限,且缺乏显式条件输入以约束工具-动作特定的安全交互区域。为此,我们提出AffordTissue,一种多模态框架,用于在胆囊切除术中预测工具-动作特异的组织可操作性区域,输出为密集热图。该方法结合时序视觉编码器(捕捉多视角下器械运动与组织动态)、语言条件输入(实现跨多样器械-动作对的泛化),以及基于DiT风格的解码器(完成密集可操作性预测)。我们通过整理并标注103例胆囊切除术中的15,638段视频,建立了首个组织可操作性基准,涵盖六种独特工具-动作组合(四种器械:钩子、抓钳、剪刀、夹闭器,对应分离、抓握、夹闭、切割任务)。实验表明,相较于视觉-语言模型基线(60.2像素平均表面距离),本方法达到20.6像素的显著改进,证明其任务特异性架构在密集手术可操作性预测上优于大规模基础模型。通过预测工具-动作特异的组织可操作区域,AffordTissue为安全手术自动化提供显式空间推理,有望实现针对合适组织区域的明确策略引导,并在器械偏离预测安全区时触发早期安全停止。
原文摘要 · Abstract (English)
Surgical action automation has progressed rapidly toward achieving surgeon-like dexterous control, driven primarily by advances in learning from demonstration and vision-language-action models. While these have demonstrated success in table-top experiments, translating them to clinical deployment remains challenging: current methods offer limited predictability on where instruments will interact on tissue surfaces and lack explicit conditioning inputs to enforce tool-action-specific safe interaction regions. Addressing this gap, we introduce AffordTissue, a multimodal framework for predicting tool-action specific tissue affordance regions as dense heatmaps during cholecystectomy. Our approach combines a temporal vision encoder capturing tool motion and tissue dynamics across multiple viewpoints, language conditioning enabling generalization across diverse instrument-action pairs, and a DiT-style decoder for dense affordance prediction. We establish the first tissue affordance benchmark by curating and annotating 15,638 video clips across 103 cholecystectomy procedures, covering six unique tool-action pairs involving four instruments (hook, grasper, scissors, clipper) and their associated tasks: dissection, grasping, clipping, and cutting. Experiments demonstrate substantial improvement over vision-language model baselines (20.6 px ASSD vs. 60.2 px for Molmo-VLM), showing that our task-specific architecture outperforms large-scale foundation models for dense surgical affordance prediction. By predicting tool-action specific tissue affordance regions, AffordTissue provides explicit spatial reasoning for safe surgical automation, potentially unlocking explicit policy guidance toward appropriate tissue regions and early safe stop when instruments deviate outside predicted safe zones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。