arXiv:2505.19306cs.ROcs.CV2025-05被引 2

用单张图片生成可避障的运动策略,靠视频建模环境结构。

From Single Images to Motion Policies via Video-Generation Environment Representations

  • 通过生成移动视角视频,从单图构建多视图数据。
  • 在真实场景中实现平滑避障运动,无需额外传感器。
  • 适合做视觉导航的机器人系统研究者参考。

自主机器人需构建环境表征并适应其几何结构以生成无碰撞运动。本文提出视频生成环境表征(VGER)框架,仅需单张输入RGB图像,即可生成符合场景几何结构的运动策略。传统方法依赖单目深度估计(如DepthAnything),但存在锥形误差。VGER利用大规模视频生成模型,以输入图像为条件生成移动相机视频,形成多视图数据集;再将帧输入预训练3D基础模型,重建稠密点云。随后采用多尺度噪声方法训练隐式环境结构表示,并构建几何一致的运动生成模型。在多样化的室内外环境中进行广泛评估,结果表明:仅凭单张图像,即可生成平滑、符合场景结构的运动轨迹。

原文摘要 · Abstract (English)

Autonomous robots typically need to construct representations of their surroundings and adapt their motions to the geometry of their environment. Here, we tackle the problem of constructing a policy model for collision-free motion generation, consistent with the environment, from a single input RGB image. Extracting 3D structures from a single image often involves monocular depth estimation. Developments in depth estimation have given rise to large pre-trained models such as DepthAnything. However, using outputs of these models for downstream motion generation is challenging due to frustum-shaped errors that arise. Instead, we propose a framework known as Video-Generation Environment Representation (VGER), which leverages the advances of large-scale video generation models to generate a moving camera video conditioned on the input image. Frames of this video, which form a multiview dataset, are then input into a pre-trained 3D foundation model to produce a dense point cloud. We then introduce a multi-scale noise approach to train an implicit representation of the environment structure and build a motion generation model that complies with the geometry of the representation. We extensively evaluate VGER over a diverse set of indoor and outdoor environments. We demonstrate its ability to produce smooth motions that account for the captured geometry of a scene, all from a single RGB input image.

视觉导航运动规划3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。