arXiv:2509.14967cs.ROcs.HC2025-09

让手术机器人理解模糊指令,靠视觉上下文和工具能力推理

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

  • 用多模态模型分析手术场景,再结合工具能力知识库推理指令
  • 在胆囊切除视频数据集上实现60%的模糊指令消歧率
  • 采用双重置信度预测保障安全,适合临床协作机器人研发

手术中人机协作受口头指令固有模糊性影响。本文提出一种机器人手术助手框架,通过将口头指令与术野视觉上下文对齐来解析和消歧。系统采用两级基于功能性的推理机制:首先利用多模态视觉-语言模型分析手术场景,再基于工具能力知识库推理解释指令。为确保患者安全,引入双集合置信度预测方法,提供统计严格的风险评估,使机器人能识别并标记模糊指令。我们在胆囊切除手术视频中构建的模糊指令数据集上验证了该框架,实现60%的通用消歧率,提出了一种更安全的人机协作方法。

原文摘要 · Abstract (English)

Effective human-robot collaboration in surgery is affected by the inherent ambiguity of verbal communication. This paper presents a framework for a robotic surgical assistant that interprets and disambiguates verbal instructions from a surgeon by grounding them in the visual context of the operating field. The system employs a two-level affordance-based reasoning process that first analyzes the surgical scene using a multimodal vision-language model and then reasons about the instruction using a knowledge base of tool capabilities. To ensure patient safety, a dual-set conformal prediction method is used to provide a statistically rigorous confidence measure for robot decisions, allowing it to identify and flag ambiguous commands. We evaluated our framework on a curated dataset of ambiguous surgical requests from cholecystectomy videos, demonstrating a general disambiguation rate of 60% and presenting a method for safer human-robot interaction in the operating room.

手术机器人指令消歧视觉语言安全推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。