从视频中学习机器人抓取与运动,兼顾可行性与避障。
Joint Flow Trajectory Optimization For Feasible Robot Motion Generation from Video Demonstrations
- 以物体为中心建模,优化抓取姿态与轨迹
- 融合抓取相似性、轨迹概率与碰撞惩罚,实现可执行路径
- 适用于真实场景下的多任务机器人操作
基于视频示教的机器人学习为替代遥操作或力导教学提供了可扩展方案,但受限于人机差异和关节可行性约束。本文提出联合流轨迹优化(JFTO)框架,在视频引导的示教学习范式下实现抓取姿态生成与物体轨迹模仿。方法不直接复制人类手部动作,而是将示范视为以物体为中心的指导,平衡三个目标:(i) 选择可行的抓取姿态,(ii) 生成与示范一致的物体轨迹,(iii) 确保在机器人运动学范围内无碰撞执行。为捕捉示范的多模态特性,将流匹配拓展至 $ ext{SE}(3)$ 空间,实现对象轨迹的概率建模,支持密度感知的模仿,避免模式崩溃。优化目标统一整合抓取相似性、轨迹似然性和碰撞惩罚,构成可微分的目标函数。我们在多种真实世界操作任务中,通过仿真与实物实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Learning from human video demonstrations offers a scalable alternative to teleoperation or kinesthetic teaching, but poses challenges for robot manipulators due to embodiment differences and joint feasibility constraints. We address this problem by proposing the Joint Flow Trajectory Optimization (JFTO) framework for grasp pose generation and object trajectory imitation under the video-based Learning-from-Demonstration (LfD) paradigm. Rather than directly imitating human hand motions, our method treats demonstrations as object-centric guides, balancing three objectives: (i) selecting a feasible grasp pose, (ii) generating object trajectories consistent with demonstrated motions, and (iii) ensuring collision-free execution within robot kinematics. To capture the multimodal nature of demonstrations, we extend flow matching to $\SE(3)$ for probabilistic modeling of object trajectories, enabling density-aware imitation that avoids mode collapse. The resulting optimization integrates grasp similarity, trajectory likelihood, and collision penalties into a unified differentiable objective. We validate our approach in both simulation and real-world experiments across diverse real-world manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。