arXiv:2606.11396cs.RO2026-06

让机械手在未知物理参数下精准操作,靠的是实时推断参数并调整动作。

PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation

论文配图:PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation
图 1 · 摘自论文原文
  • 构建联合概率模型,同时估计物体参数和动态变化。
  • 在仿真中训练后零样本迁移至真实硬件,成功率超现有方法。
  • 适合高精度抓取、拧螺丝等需感知物理特性的任务场景。

多指灵巧操作对物体形状、姿态和摩擦系数等物理参数敏感。尽管仿真可提供大量带已知参数的数据,但部署时真实参数未知,导致模型性能下降。传统领域随机化策略难以应对如拧螺丝等精确任务,因策略需随具体参数调整。为此,我们提出概率潜在统一世界建模与参数估计(PLUME),该模型联合学习参数信念演化及条件于参数的系统动力学。通过潜空间共同表征多种物理参数与奖励函数(依赖部分可观测变量),以支持规划。新颖的学习框架实现在线参数推断,无需重训练或微调即可高效对齐真实动态。我们在模拟的螺丝刀旋转、阀门转动、桶提拉和碟片拨动任务以及真实硬件螺丝刀任务上评估,结果表明:仿真训练策略可成功实现零样本迁移,且优于最先进的离线强化学习和世界模型增强行为克隆基线。

原文摘要 · Abstract (English)

Dexterous manipulation with multi-finger hands can be sensitive to physical parameters such as object shape, pose, and friction coefficients. While simulation enables large-scale data collection with known parameter values, simulation-trained policies must still handle uncertainty at deployment, where the true parameters and therefore the true dynamics are unknown. Standard domain randomization strategies may be insufficient for precise tasks like screwdriver turning, as manipulation strategies may need to change depending on specific parameter values. To address this, we propose Probabilistic Latent Unified world Modeling and parameter Estimation (PLUME), a world model that jointly learns to evolve a belief over parameter values as well as the system dynamics conditioned on those parameters. We learn a latent space to jointly represent multiple qualitatively different physical parameters along with rewards, themselves functions of partially-observable variables, to inform planning. Our novel learning framework leads to efficient alignment of the world model to true dynamics through online parameter inference as opposed to re-training or fine-tuning. We evaluate our method on simulated screwdriver turning, valve turning, bucket lifting, and disk flicking tasks, as well as a hardware screwdriver turning task, where we achieve successful zero-shot transfer of our simulation-trained policy and outperform state-of-the-art offline reinforcement learning and world-model-augmented behavior cloning baselines. Please see our website at https://plume-world-model.github.io for videos.

灵巧操作世界模型参数估计零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。