arXiv:2511.07392cs.CLcs.AI2025-11

让外科医生用语音操控手术数据,不中断操作流程。

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

  • 基于大模型的多智能体框架,分层处理语音指令。
  • 在240条指令上实现高准确率与强容错能力。
  • 适合希望提升手术效率的医疗科技团队。

在机器人手术中,外科医生需全程专注双手和视觉操作,难以中断流程获取或操作多模态患者数据。为此,我们提出语音交互式手术代理(VISA),采用分层多智能体架构,包含一个协调代理和三个由大语言模型驱动的任务专用代理。这些代理能自主规划、优化、验证与推理,理解语音指令并执行如调取临床信息、操作CT影像或导航3D解剖模型等任务。我们构建了包含240条用户指令的层级化数据集,并提出多级协调评估指标(MOEM),从命令与类别两个层面评估性能与鲁棒性。实验表明,VISA在阶段级与流程级均表现出高准确率与成功率,且具备纠正转录错误、化解语言歧义、解析多样化自由表达的能力。结果凸显其在机器人手术中的潜力,以及未来集成新功能与代理的可扩展性。

原文摘要 · Abstract (English)

In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, we propose a Voice-Interactive Surgical Agent (VISA) built on a hierarchical multi-agent framework consisting of an orchestration agent and three task-specific agents driven by Large Language Models (LLMs). These LLM-based agents autonomously plan, refine, validate, and reason to interpret voice commands and execute tasks such as retrieving clinical information, manipulating CT scans, or navigating 3D anatomical models within surgical video. We construct a dataset of 240 user commands organized into hierarchical categories and introduce the Multi-level Orchestration Evaluation Metric (MOEM) that evaluates the performance and robustness at both the command and category levels. Experimental results demonstrate that VISA achieves high stage-level accuracy and workflow-level success rates, while also enhancing its robustness by correcting transcription errors, resolving linguistic ambiguity, and interpreting diverse free-form expressions. These findings highlight the strong potential of VISA to support robotic surgery and its scalability for integrating new functions and agents.

手术机器人语音控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。