arXiv:2601.08454cs.RO2026-01被引 1

用大模型自动规划机器人动作,高效构建真实物体的物理仿真

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs

  • 用视觉语言模型解析任务意图,生成可执行的机器人行为树
  • 实测仅需少量接触交互,就准确估计物体质量、摩擦等关键参数
  • 自动生成的行为树兼具安全性和可解释性,适合工业数字孪生应用

传统的真实世界到仿真环境构建(Real2Sim)依赖人工系统辨识或盲目探索,缺乏语义上下文理解,导致冗余交互与低效数据采集。本文提出一种自主、目标驱动的Real2Sim框架,利用视觉语言模型(VLMs)进行语义任务分解。给定自然语言指令、不完整的仿真描述和视觉观测,该框架能自动识别完成任务所需的最小缺失物理参数集,并生成由原子运动与感知原语构成的反应式行为树(BT),通过丰富的接触式机器人交互选择性获取这些参数。在扭矩控制的Franka Emika Panda机器人上开展的大量真实实验表明,该方法可准确估计物体质量、表面几何及摩擦系数等衍生参数。定量评估显示,相比全面探索基线方法,操作效率显著提升;消融实验验证了提示架构在不同主流VLM上的鲁棒性。此外,行为树的层次结构作为确定性安全过滤器,有效抑制了生成式VLM幻觉,防止了不安全物理异常。本工作为从非结构化人类意图直接构建物理感知数字孪生提供了一条可扩展、高效且可解释的路径。

原文摘要 · Abstract (English)

Constructing physically accurate simulation environments (Real2Sim) traditionally relies on manual system identification or rigid, exhaustive exploration routines. These task-agnostic pipelines often fail to leverage semantic scene context, leading to redundant physical interactions and inefficient data acquisition. In this paper, we present an autonomous, intent-driven Real2Sim framework that leverages Vision-Language Models (VLMs) for Semantic Task Decomposition. Given a high-level natural language request, an incomplete simulation description, and a visual observation, the framework autonomously identifies the minimal subset of missing physical parameters required for the simulation task. It then generates a reactive Behavior Tree (BT) composed of atomic motion and sensing primitives to selectively acquire these parameters through contact-rich robotic interaction. Extensive real-world experiments on a torque-controlled Franka Emika Panda demonstrate that our approach accurately estimates object mass, surface geometry, and derived parameters such as friction. Quantitative evaluations reveal significant operational efficiency gains compared to exhaustive baseline methods, while ablation studies confirm the robustness of the prompt architecture across different state-of-the-art VLMs. Furthermore, the reactive hierarchy of the BT acts as a deterministic safety filter, successfully mitigating generative VLM hallucinations and preventing unsafe physical anomalies. Ultimately, this work provides a scalable, efficient, and interpretable pipeline for building physics-aware digital twins directly from unstructured human intent.

数字孪生机器人学习大模型仿真建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。