用语义数字孪生让大模型更懂环境,机器人能自适应完成家务任务。
Grounding Language Models with Semantic Digital Twins for Robotic Planning
- 将自然语言转为带上下文的动作三元组,由数字孪生提供环境语义支持。
- 在任务失败时结合错误反馈与环境数据,自动生成修复方案并重规划。
- 适用于家庭场景的复杂任务,尤其适合需要动态适应的机器人系统。
我们提出一种新框架,将语义数字孪生(SDTs)与大语言模型(LLMs)结合,实现动态环境中自适应、目标驱动的机器人任务执行。系统将自然语言指令分解为结构化的动作三元组,并基于SDT提供的上下文环境数据进行语义定位,使机器人能够理解物体功能与交互规则,从而实现动作规划与实时调整。当执行失败时,LLM利用错误反馈和SDT信息生成恢复策略,并迭代修正动作计划。我们在ALFRED基准的任务上进行了评估,展示了在多种家庭场景下的鲁棒表现。该框架有效融合高层推理与语义环境理解,在不确定性与故障情况下仍能可靠完成任务。
原文摘要 · Abstract (English)
We introduce a novel framework that integrates Semantic Digital Twins (SDTs) with Large Language Models (LLMs) to enable adaptive and goal-driven robotic task execution in dynamic environments. The system decomposes natural language instructions into structured action triplets, which are grounded in contextual environmental data provided by the SDT. This semantic grounding allows the robot to interpret object affordances and interaction rules, enabling action planning and real-time adaptability. In case of execution failures, the LLM utilizes error feedback and SDT insights to generate recovery strategies and iteratively revise the action plan. We evaluate our approach using tasks from the ALFRED benchmark, demonstrating robust performance across various household scenarios. The proposed framework effectively combines high-level reasoning with semantic environment understanding, achieving reliable task completion in the face of uncertainty and failure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。