arXiv:2601.18121cs.ROcs.CV2026-01

将视觉对齐的手物交互数据转为物理可执行轨迹,提升仿真稳定性。

Grasp-and-Lift: Executable 3D Hand-Object Interaction Reconstruction via Physics-in-the-Loop Optimization

  • 用稀疏关键帧的样条参数化手部运动,构建黑箱优化问题。
  • 在物理仿真中实现稳定抓取与举起,手物姿态误差低于现有方法。
  • 适合需要真实物理交互数据的机器人抓取策略训练场景。

灵巧手操作越来越依赖大规模、高精度手物轨迹数据集。然而,当前主流数据集如DexYCB和HO3D主要针对视觉对齐优化,在物理仿真中重放时常出现穿透、接触丢失和抓取不稳等物理不可行问题。本文提出一种仿真闭环优化框架,将视觉对齐轨迹转化为物理可执行轨迹。核心创新在于将该问题建模为可求解的黑箱优化问题:采用基于稀疏时间关键帧的低维样条表示手部运动,利用强大的无梯度优化器CMA-ES,将高保真物理引擎作为黑箱目标函数。该方法同时最大化物理成功性(如稳定抓取与举起)并最小化与原始人类示范的偏差。相比MANIPTRANS等近期迁移管道,本方法在重放过程中显著降低手部与物体的姿态误差,更准确恢复手物物理交互。本方法提供了一种通用且可扩展的方案,用于将视觉演示转换为物理有效轨迹,支持生成高质量数据以训练鲁棒控制策略。

原文摘要 · Abstract (English)

Dexterous hand manipulation increasingly relies on large-scale motion datasets with precise hand-object trajectory data. However, existing resources such as DexYCB and HO3D are primarily optimized for visual alignment but often yield physically implausible interactions when replayed in physics simulators, including penetration, missed contact, and unstable grasps. We propose a simulation-in-the-loop refinement framework that converts these visually aligned trajectories into physically executable ones. Our core contribution is to formulate this as a tractable black-box optimization problem. We parameterize the hand's motion using a low-dimensional, spline-based representation built on sparse temporal keyframes. This allows us to use a powerful gradient-free optimizer, CMA-ES, to treat the high-fidelity physics engine as a black-box objective function. Our method finds motions that simultaneously maximize physical success (e.g., stable grasp and lift) while minimizing deviation from the original human demonstration. Compared to MANIPTRANS-recent transfer pipelines, our approach achieves lower hand and object pose errors during replay and more accurately recovers hand-object physical interactions. Our approach provides a general and scalable method for converting visual demonstrations into physically valid trajectories, enabling the generation of high-fidelity data crucial for robust policy learning.

手物交互物理仿真轨迹优化机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。