arXiv:2505.10359cs.ROcs.CV2025-05被引 4

通过自适应生成新视角图像,提升机器人在复杂环境中的语言指令执行能力。

NVSPolicy: Adaptive Novel-View Synthesis for Generalizable Language-Conditioned Policy Learning

论文配图:NVSPolicy: Adaptive Novel-View Synthesis for Generalizable Language-Conditioned Policy Learning
图 1 · 摘自论文原文
  • 动态选择视角并生成新视图以补充视觉信息
  • 在CALVIN数据集上实现90.4%的平均成功率,显著优于现有方法
  • 适合需要强泛化能力的现实机器人任务,尤其适用于语言控制场景

近期深度生成模型展现出前所未有的零样本泛化能力,为非结构化环境中机器人操作提供了巨大潜力。给定场景的部分观测,生成模型可补全未见区域,从而增强机器人对未知环境的泛化能力。然而,由于生成图像存在视觉伪影以及多模态特征融合效率低,该方向仍面临挑战。本文提出NVSPolicy,一种耦合自适应新视角合成模块与分层策略网络的语言条件策略学习方法。输入图像后,NVSPolicy动态选择有信息量的视角,并生成自适应新视角图像以丰富视觉上下文。为缓解生成图像不完美带来的影响,采用循环一致性变分自编码器机制,将视觉特征解耦为语义特征与剩余特征。二者分别输入分层策略网络:语义特征用于高层元技能选择,剩余特征指导低层动作估计。此外,设计多种实用机制提升方法效率。在CALVIN上的大量实验表明,本方法达到90.4%的平均成功率,显著超越近期方法。消融实验证明了自适应新视角合成范式的有效性。进一步在真实机器人平台上的评估展示了其实际应用价值。

原文摘要 · Abstract (English)

Recent advances in deep generative models demonstrate unprecedented zero-shot generalization capabilities, offering great potential for robot manipulation in unstructured environments. Given a partial observation of a scene, deep generative models could generate the unseen regions and therefore provide more context, which enhances the capability of robots to generalize across unseen environments. However, due to the visual artifacts in generated images and inefficient integration of multi-modal features in policy learning, this direction remains an open challenge. We introduce NVSPolicy, a generalizable language-conditioned policy learning method that couples an adaptive novel-view synthesis module with a hierarchical policy network. Given an input image, NVSPolicy dynamically selects an informative viewpoint and synthesizes an adaptive novel-view image to enrich the visual context. To mitigate the impact of the imperfect synthesized images, we adopt a cycle-consistent VAE mechanism that disentangles the visual features into the semantic feature and the remaining feature. The two features are then fed into the hierarchical policy network respectively: the semantic feature informs the high-level meta-skill selection, and the remaining feature guides low-level action estimation. Moreover, we propose several practical mechanisms to make the proposed method efficient. Extensive experiments on CALVIN demonstrate the state-of-the-art performance of our method. Specifically, it achieves an average success rate of 90.4\% across all tasks, greatly outperforming the recent methods. Ablation studies confirm the significance of our adaptive novel-view synthesis paradigm. In addition, we evaluate NVSPolicy on a real-world robotic platform to demonstrate its practical applicability.

机器人学习生成模型语言控制新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。