arXiv:2603.14393cs.RO2026-03被引 1

用大模型让机器人自主读指南、调工具,实现多部位超声自动扫描。

From Scanning Guidelines to Action: A Robotic Ultrasound Agent with LLM-Based Reasoning

  • 大模型解析超声扫描指南,动态调用工具执行任务。
  • 在胆囊、脊柱、肾脏上实现可靠自主扫描,适应复杂决策流程。
  • 通过强化学习提升推理与工具选择准确率,适合医疗机器人研发者。

机器人超声相比手动扫描具有更好的可重复性和更低的操作依赖性。临床中,超声采集高度依赖超声医师的经验与情境判断。将该过程移植到机器人系统时,常通过固定流程和专用模型显式编码经验,导致难以适应新任务。本文提出一种统一的自主机器人超声扫描框架,利用大模型代理解析超声扫描指南,并动态调用预设软件工具执行扫描。该代理不依赖固定流程,而是从操作手册中检索并推理指南步骤,根据观察结果与当前扫描状态调整规划。系统可处理变量性与决策依赖型工作流,如调整策略、重复步骤或根据图像质量或解剖发现选择下一工具调用。为提升工具选择背后的推理透明度与可信度,进一步采用基于强化学习的方法微调大模型代理,同时提升推理质量与工具选择及参数设置的正确性,保持对未见指南及相关任务的强泛化能力。首先通过10个超声扫描指南的口头执行验证方法,评估推理及工具调用与参数化表现,证明强化学习微调的有效性。随后在胆囊、脊柱、肾脏的机器人扫描中展示真实世界可行性。整体框架能遵循多样指南,在统一系统中实现多个解剖目标的可靠自主扫描。

原文摘要 · Abstract (English)

Robotic ultrasound offers advantages over free-hand scanning, including improved reproducibility and reduced operator dependency. In clinical practice, US acquisition relies heavily on the sonographer's experience and situational judgment. When transferring this process to robotic systems, such expertise is often encoded explicitly through fixed procedures and task-specific models, yielding pipelines that can be difficult to adapt to new scanning tasks. In this work, we propose a unified framework for autonomous robotic US scanning that leverages a LLM-based agent to interpret US scanning guidelines and execute scans by dynamically invoking a set of provided software tools. Instead of encoding fixed scanning procedures, the LLM agent retrieves and reasons over guideline steps from scanning handbooks and adapts its planning decisions based on observations and the current scanning state. This enables the system to handle variable and decision-dependent workflows, such as adjusting scanning strategies, repeating steps, or selecting the appropriate next tool call in response to image quality or anatomical findings. Because the reasoning underlying tool selection is also critical for transparent and trustworthy planning, we further fine tune the LLM agent using a RL based strategy to improve both its reasoning quality and the correctness of tool selection and parameterization, while maintaining robust generalization to unseen guidelines and related tasks. We first validate the approach via verbal execution on 10 US scanning guidelines, assessing reasoning as well as tool selection and parameterization, and showing the benefit of RL fine tuning. We then demonstrate real world feasibility on robotic scanning of the gallbladder, spine, and kidney. Overall, the framework follows diverse guidelines and enables reliable autonomous scanning across multiple anatomical targets within a unified system.

机器人超声大模型应用自主导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。