arXiv:2507.11525cs.ROcs.HC2025-07中稿 · 2025 IEEE Internat…被引 5

用大模型检测手术指令歧义,提升人机协作安全性。

LLM-based ambiguity detection in natural language instructions for collaborative surgical robots

  • 用多个提示策略的LLM组合检测语言、上下文等四类歧义
  • 在手术指令上实现超60%的歧义分类准确率
  • 适合医疗机器人安全验证与智能辅助系统开发

自然语言指令中的歧义在安全关键型人机交互中带来显著风险,尤其在手术领域。为此,我们提出一个专为协同手术场景设计的基于大语言模型(LLM)的歧义检测框架。该方法采用一组配备不同提示策略的LLM评估器,分别识别语言、上下文、流程和关键性歧义;引入链式思维评估器系统分析指令结构潜在问题。各评估结果通过置信区间预测融合,基于与标注校准数据集的比较生成非符合度分数。在Llama 3.2 11B和Gemma 3 12B上测试,对歧义与非歧义手术指令的分类准确率超过60%。该方法通过在机器人执行前识别潜在歧义指令,提升了手术中人机协作的安全性与可靠性。

原文摘要 · Abstract (English)

Ambiguity in natural language instructions poses significant risks in safety-critical human-robot interaction, particularly in domains such as surgery. To address this, we propose a framework that uses Large Language Models (LLMs) for ambiguity detection specifically designed for collaborative surgical scenarios. Our method employs an ensemble of LLM evaluators, each configured with distinct prompting techniques to identify linguistic, contextual, procedural, and critical ambiguities. A chain-of-thought evaluator is included to systematically analyze instruction structure for potential issues. Individual evaluator assessments are synthesized through conformal prediction, which yields non-conformity scores based on comparison to a labeled calibration dataset. Evaluating Llama 3.2 11B and Gemma 3 12B, we observed classification accuracy exceeding 60% in differentiating ambiguous from unambiguous surgical instructions. Our approach improves the safety and reliability of human-robot collaboration in surgery by offering a mechanism to identify potentially ambiguous instructions before robot action.

大模型手术机器人歧义检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。