arXiv:2511.22815cs.CV2025-11被引 4

用姿态对齐的3D记忆实现精准长视频生成与相机控制。

Captain Safari: A World Engine with Pose-Aligned 3D Memory

  • 通过姿态条件检索3D世界记忆,动态生成视频。
  • 在复杂路径下误差降低至0.3690,轨迹跟随提升至0.200。
  • 适合需要高精度3D一致性的交互式视频生成场景。

世界引擎旨在生成支持用户控制摄像机运动的长时、3D一致视频。然而,现有系统在激进的6自由度轨迹和复杂户外布局下表现不佳:长期几何一致性丢失、偏离目标路径或退化为过度保守的运动。为此,我们提出Captain Safari,一种基于姿态条件的世界引擎,通过从持久的世界记忆中检索姿态对齐的世界标记来生成视频。给定相机路径,该方法维护动态局部记忆,并使用检索器获取姿态对齐的世界标记,进而引导沿轨迹的视频生成。这一设计使模型在执行复杂摄像机动画的同时保持稳定的3D结构。为评估该设定,我们构建了OpenSafari——一个全新的野外第一人称视角数据集,包含经多阶段几何与运动验证的高动态无人机视频。在视频质量、3D一致性和轨迹跟随性方面,Captain Safari显著优于最先进相机控制生成器:MEt3R从0.3703降至0.3690,AUC@30从0.181提升至0.200,且FVD远低于所有基线。更重要的是,在50名参与者、五模型对比的人类评估中,67.6%的选择偏好本方法。结果表明,姿态条件世界记忆是长时程可控视频生成的强大机制,并提供OpenSafari作为未来世界引擎研究的新基准。

原文摘要 · Abstract (English)

World engines aim to synthesize long, 3D-consistent videos that support interactive exploration of a scene under user-controlled camera motion. However, existing systems struggle under aggressive 6-DoF trajectories and complex outdoor layouts: they lose long-range geometric coherence, deviate from the target path, or collapse into overly conservative motion. To this end, we introduce Captain Safari, a pose-conditioned world engine that generates videos by retrieving from a persistent world memory. Given a camera path, our method maintains a dynamic local memory and uses a retriever to fetch pose-aligned world tokens, which then condition video generation along the trajectory. This design enables the model to maintain stable 3D structure while accurately executing challenging camera maneuvers. To evaluate this setting, we curate OpenSafari, a new in-the-wild FPV dataset containing high-dynamic drone videos with verified camera trajectories, constructed through a multi-stage geometric and kinematic validation pipeline. Across video quality, 3D consistency, and trajectory following, Captain Safari substantially outperforms state-of-the-art camera-controlled generators. It reduces MEt3R from 0.3703 to 0.3690, improves AUC@30 from 0.181 to 0.200, and yields substantially lower FVD than all camera-controlled baselines. More importantly, in a 50-participant, 5-way human study where annotators select the best result among five anonymized models, 67.6% of preferences favor our method across all axes. Our results demonstrate that pose-conditioned world memory is a powerful mechanism for long-horizon, controllable video generation and provide OpenSafari as a challenging new benchmark for future world-engine research.

3D生成视频合成世界引擎姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。