arXiv:2505.01729cs.CV2025-05中稿 · IEEE/RSJ IROS 2025被引 4

用自监督深度控制相机位姿,提升生成世界模型的视角一致性。

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth

  • 通过自监督深度与运动估计,实现视频帧间位姿联动。
  • 在自动驾驶与通用视频数据集上显著提升视角合成精度。
  • 适合需要精准视角控制的生成式世界模型研究者。

自动驾驶系统近年来的发展凸显了世界模型在常规与复杂驾驶场景中实现鲁棒、泛化性能的潜力。然而,精确且灵活的相机位姿控制仍是关键挑战,这对准确的视角变换和场景动态的真实模拟至关重要。本文提出PosePilot,一种轻量但强大的框架,显著增强生成式世界模型中的相机位姿可控性。受自监督深度估计启发,PosePilot采用结构光从运动(SfM)原理,建立相机位姿与视频生成之间的紧密耦合。具体而言,引入自监督深度与位姿读出,使模型能直接从视频序列中推断深度和相对相机运动。这些输出驱动基于光度一致性损失的位姿感知帧扭曲,确保合成帧间的几何一致性。为进一步优化位姿估计,引入逆向扭曲步骤与位姿回归损失,提升视角精度与适应性。在自动驾驶及通用视频数据集上的大量实验表明,PosePilot显著增强了扩散模型与自回归模型的结构理解与运动推理能力。通过自监督深度引导相机位姿调控,PosePilot为位姿可控性树立新基准,实现了物理一致、可靠的视角合成。

原文摘要 · Abstract (English)

Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge remains: precise and flexible camera pose control, which is crucial for accurate viewpoint transformation and realistic simulation of scene dynamics. In this paper, we introduce PosePilot, a lightweight yet powerful framework that significantly enhances camera pose controllability in generative world models. Drawing inspiration from self-supervised depth estimation, PosePilot leverages structure-from-motion principles to establish a tight coupling between camera pose and video generation. Specifically, we incorporate self-supervised depth and pose readouts, allowing the model to infer depth and relative camera motion directly from video sequences. These outputs drive pose-aware frame warping, guided by a photometric warping loss that enforces geometric consistency across synthesized frames. To further refine camera pose estimation, we introduce a reverse warping step and a pose regression loss, improving viewpoint precision and adaptability. Extensive experiments on autonomous driving and general-domain video datasets demonstrate that PosePilot significantly enhances structural understanding and motion reasoning in both diffusion-based and auto-regressive world models. By steering camera pose with self-supervised depth, PosePilot sets a new benchmark for pose controllability, enabling physically consistent, reliable viewpoint synthesis in generative world models.

世界模型相机位姿自监督视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。