用视觉大模型构建手术数字孪生,提升机器人手术的鲁棒性
Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models
- 基于视觉大模型生成手术环境的数字孪生表示
- 在dVRK平台上实现钉板传递与纱布抓取任务,表现稳定
- 适合关注手术自动化与具身智能融合的研究者
基于大语言模型(LLM)的智能体正成为实现鲁棒具身智能的强大工具,因其具备规划复杂动作序列的能力。在手术自动化中,这种规划能力尤为关键。然而,这类智能体依赖于场景的详细自然语言描述,因此需要强大且鲁棒的感知算法,从视觉输入中提取精细的场景表征。此前研究多聚焦于基于LLM的任务规划,而采用简单但严重受限的感知方案,难以拓展至非受控环境。本文提出一种基于数字孪生(DT)的机器感知新方法,利用近期视觉基础模型的出色性能与零样本泛化能力。将该数字孪生表示与LLM规划智能体结合,并部署于dVRK平台,构建具身智能系统,在钉板传递与纱布抓取任务中验证其鲁棒性。结果表明,该方法具备强任务表现和对多样化环境的泛化能力。尽管表现优异,本工作仍为数字孪生集成的初步探索,未来需进一步构建完整框架以提升手术中具身智能的可解释性与泛化性。
原文摘要 · Abstract (English)
Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning complex action sequences. Sound planning ability is necessary for robust automation in many task domains, but especially in surgical automation. These agents rely on a highly detailed natural language representation of the scene. Thus, to leverage the emergent capabilities of LLM agents for surgical task planning, developing similarly powerful and robust perception algorithms is necessary to derive a detailed scene representation of the environment from visual input. Previous research has focused primarily on enabling LLM-based task planning while adopting simple yet severely limited perception solutions to meet the needs for bench-top experiments, but lacks the critical flexibility to scale to less constrained settings. In this work, we propose an alternate perception approach -- a digital twin (DT)-based machine perception approach that capitalizes on the convincing performance and out-of-the-box generalization of recent vision foundation models. Integrating our DT representation and LLM agent for planning with the dVRK platform, we develop an embodied intelligence system and evaluate its robustness in performing peg transfer and gauze retrieval tasks. Our approach shows strong task performance and generalizability to varied environmental settings. Despite a convincing performance, this work is merely a first step towards the integration of DT representations. Future studies are necessary for the realization of a comprehensive DT framework to improve the interpretability and generalizability of embodied intelligence in surgery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。