arXiv:2506.01943cs.CV2025-06被引 36

通过分阶段建模物体交互,提升机器人操作视频生成的准确性。

Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control

  • 将交互过程分为三阶段,分别以机械臂和目标物为主导建模。
  • 在Bridge数据集上实现新最好效果,视觉保真度显著提升。
  • 适合研究机器人视觉生成与多物体交互建模的学者。

视频扩散模型在生成机器人决策数据方面展现出潜力,轨迹条件可实现精细控制。然而,现有方法主要关注单个物体运动,难以捕捉复杂操作中的多物体交互,原因在于重叠区域特征耦合导致视觉质量下降。为此,我们提出RoboMaster框架,通过协作轨迹建模方式刻画物体间动态关系。不同于以往将物体分离处理的方法,本工作将交互过程分解为预交互、交互和后交互三个子阶段,并在每个阶段以主导物体(预/后阶段为机械臂,交互阶段为目标物)为核心进行建模,有效缓解了多物体特征融合问题。为进一步保证视频中物体语义一致性,引入外观与形状感知的潜在表征。在挑战性数据集Bridge以及RLBench和SIMPLER基准上的大量实验表明,该方法在轨迹控制的机器人操作视频生成任务中达到新的最优性能。

原文摘要 · Abstract (English)

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and struggle to capture multi-object interaction crucial in complex manipulation. This limitation arises from entangled features in overlapping regions, leading to degraded visual fidelity. To address this, we present RoboMaster, a novel framework that models inter-object dynamics via a collaborative trajectory formulation. Unlike prior methods that decompose objects, our core is to decompose the interaction process into three sub-stages: pre-interaction, interaction, and post-interaction, and models each phase using the dominant object, specifically the robotic arm in the pre- and post-interaction phases and the manipulated object during interaction. This design effectively alleviates the multi-object feature fusion issue in prior work. To further ensure subject semantic consistency across the video, we incorporate appearance- and shape-aware latent representations for objects. Extensive experiments on the challenging Bridge dataset, as well as RLBench and SIMPLER benchmarks, demonstrate that our method establishs new state-of-the-art performance in trajectory-controlled video generation for robotic manipulation. Project Page: https://fuxiao0719.github.io/projects/robomaster/

视频生成机器人操作扩散模型多物体交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。