arXiv:2603.20992cs.RO2026-03

用物理仿真优化物体姿态,让机械手抓取更真实可靠。

Geometrically Plausible Object Pose Refinement using Differentiable Simulation

  • 结合物理仿真与多模态感知,优化物体姿态的合理性
  • 初始误差高时,手物相交体积减少超87%
  • 适合需要精准抓取的机器人操作场景

当前最先进的物体姿态估计方法容易产生几何上不可行的姿态假设,这在灵巧操作中尤为突出,例如估计姿态常与机械手发生穿插或脱离支撑面。本文提出一种多模态姿态精修方法,融合可微分物理仿真、可微分渲染和视觉触觉传感,同时优化姿态的空间精度与物理一致性。模拟实验表明,当初始估计准确时,物体与机械手之间的相交体积误差降低73%;在高初始不确定性下,该误差降低超过87%,显著优于传统的ICP基线方法。此外,几何合理性提升的同时,平移与旋转误差也同步下降。实现既符合物理现实又忠实于多模态传感器输入的姿态估计,是迈向鲁棒手内操作的关键一步。

原文摘要 · Abstract (English)

State-of-the-art object pose estimation methods are prone to generating geometrically infeasible pose hypotheses. This problem is prevalent in dexterous manipulation, where estimated poses often intersect with the robotic hand or are not lying on a support surface. We propose a multi-modal pose refinement approach that combines differentiable physics simulation, differentiable rendering and visuo-tactile sensing to optimize object poses for both spatial accuracy and physical consistency. Simulated experiments show that our approach reduces the intersection volume error between the object and robotic hand by 73\% when the initial estimate is accurate and by over 87\% under high initial uncertainty, significantly outperforming standard ICP-based baselines. Furthermore, the improvement in geometric plausibility is accompanied by a concurrent reduction in translation and orientation errors. Achieving pose estimation that is grounded in physical reality while remaining faithful to multi-modal sensor inputs is a critical step toward robust in-hand manipulation.

姿态估计物理仿真机器人抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。