arXiv:2504.16464cs.ROcs.AI2025-04被引 12

用动作树和视觉引导提升机器人操作视频生成的准确性与质量

ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance

  • 将指令构建成动作树,通过节点嵌入捕捉指令间关系
  • 融合深度与语义引导,显著提升视频时空一致性和视觉质量
  • 在未见任务中视频质量指标全面领先,适用于复杂操作仿真

尽管近期机器人操作视频合成取得进展,但在指令遵循性和视觉质量方面仍面临挑战。现有方法如RoboDreamer虽采用语言分解将指令拆分为低层原语,并以此条件化世界模型以实现组合式指令遵循,但这些原语未考虑彼此间的关系。此外,现有方法忽视了深度与语义等关键视觉引导信息,而这些对提升视觉质量至关重要。本文提出ManipDreamer,一种基于动作树与视觉引导的先进世界模型。为更好学习指令原语间的关联,我们以动作树形式表示指令,并为每个节点分配嵌入,使每条指令可通过遍历动作树获取嵌入表示,进而指导世界模型。为增强视觉质量,引入可兼容世界模型的视觉引导适配器,融合深度与语义引导,提升视频生成的时序与物理一致性。基于上述机制,ManipDreamer显著提升指令遵循能力与视觉质量。在机器人操作基准上的综合评估显示,相比近期的RoboDreamer模型,其在已见与未见任务中均实现大幅改进:未见任务中PSNR从19.55提升至21.05,SSIM从0.7474提升至0.7982,光流误差从3.506降至3.201;同时,在6个RLbench任务上平均提升2.5%的操作成功率。

原文摘要 · Abstract (English)

While recent advancements in robotic manipulation video synthesis have shown promise, significant challenges persist in ensuring effective instruction-following and achieving high visual quality. Recent methods, like RoboDreamer, utilize linguistic decomposition to divide instructions into separate lower-level primitives, conditioning the world model on these primitives to achieve compositional instruction-following. However, these separate primitives do not consider the relationships that exist between them. Furthermore, recent methods neglect valuable visual guidance, including depth and semantic guidance, both crucial for enhancing visual quality. This paper introduces ManipDreamer, an advanced world model based on the action tree and visual guidance. To better learn the relationships between instruction primitives, we represent the instruction as the action tree and assign embeddings to tree nodes, each instruction can acquire its embeddings by navigating through the action tree. The instruction embeddings can be used to guide the world model. To enhance visual quality, we combine depth and semantic guidance by introducing a visual guidance adapter compatible with the world model. This visual adapter enhances both the temporal and physical consistency of video generation. Based on the action tree and visual guidance, ManipDreamer significantly boosts the instruction-following ability and visual quality. Comprehensive evaluations on robotic manipulation benchmarks reveal that ManipDreamer achieves large improvements in video quality metrics in both seen and unseen tasks, with PSNR improved from 19.55 to 21.05, SSIM improved from 0.7474 to 0.7982 and reduced Flow Error from 3.506 to 3.201 in unseen tasks, compared to the recent RoboDreamer model. Additionally, our method increases the success rate of robotic manipulation tasks by 2.5% in 6 RLbench tasks on average.

机器人操作视频生成动作树视觉引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。