用视觉语言模型和实时交互提升仿真到现实的机械臂操作成功率
Phys2Real: Fusing VLM Priors with Interactive Online Adaptation for Uncertainty-Aware Sim-to-Real Manipulation
- 结合视觉语言模型先验与在线交互数据融合估计物理参数
- 在不同重心的推块任务中成功率达100%且完成速度提升15%
- 适合需要精准动力学建模的机器人实操场景
直接在真实世界中训练机器人操作策略成本高且耗时。虽然在仿真中训练强化学习(RL)策略更具可扩展性,但有效的仿真到现实迁移仍具挑战,尤其对需精确动力学的任务。为此,我们提出 Phys2Real,一种从真实到仿真再到真实的强化学习流程,融合视觉语言模型(VLM)推断的物理参数先验与不确定性感知的在线交互适应。该方法包含三个核心部分:(1) 基于3D高斯点阵的高保真几何重建,(2) 由VLM推断的物理参数先验分布,(3) 通过交互数据进行的在线物理参数估计。Phys2Real以可解释的物理参数为条件,利用基于集成的不确定性量化来精炼VLM预测。在具有不同质心的T型块平面推移任务及质量分布偏心的锤子推移任务中,相较于领域随机化基线,Phys2Real分别实现100%与79%的成功率(底部加重的T块)、57%与23%的成功率(顶部加重的T块),以及锤子推移平均任务完成时间缩短15%。消融实验表明,VLM信息与交互信息的结合对成功至关重要。项目主页:https://phys2real.github.io/。
原文摘要 · Abstract (English)
Learning robotic manipulation policies directly in the real world can be expensive and time-consuming. While reinforcement learning (RL) policies trained in simulation present a scalable alternative, effective sim-to-real transfer remains challenging, particularly for tasks that require precise dynamics. To address this, we propose Phys2Real, a real-to-sim-to-real RL pipeline that combines vision-language model (VLM)-inferred physical parameter estimates with interactive adaptation through uncertainty-aware fusion. Our approach consists of three core components: (1) high-fidelity geometric reconstruction with 3D Gaussian splatting, (2) VLM-inferred prior distributions over physical parameters, and (3) online physical parameter estimation from interaction data. Phys2Real conditions policies on interpretable physical parameters, refining VLM predictions with online estimates via ensemble-based uncertainty quantification. On planar pushing tasks of a T-block with varying center of mass (CoM) and a hammer with an off-center mass distribution, Phys2Real achieves substantial improvements over a domain randomization baseline: 100% vs 79% success rate for the bottom-weighted T-block, 57% vs 23% in the challenging top-weighted T-block, and 15% faster average task completion for hammer pushing. Ablation studies indicate that the combination of VLM and interaction information is essential for success. Project website: https://phys2real.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。