arXiv:2603.17375cs.CV2026-03

用单目图像生成立体视频,直接从视差学习几何结构。

Stereo World Model: Camera-Guided Stereo Video Generation

  • 用相机感知的旋转位置编码统一建模视角与时间一致性。
  • 相比单目转立体方案,视点一致性提升5%,生成速度超3倍。
  • 适合虚拟现实、机器人交互等需真实立体感的应用场景。

我们提出 StereoWorld,一种基于相机条件的立体世界模型,联合学习外观与双目几何,实现端到端立体视频生成。不同于单目RGB或RGBD方法,StereoWorld仅使用RGB模态,直接从视差中获取几何信息。为高效实现一致的立体生成,提出两项关键设计:(1) 统一的相机帧RoPE,将相机感知的旋转位置编码注入潜在标记,实现相对视角与时间一致性,同时通过稳定的注意力初始化保留预训练视频先验;(2) 立体感知注意力分解,将全4D注意力拆分为3D视图内注意力与水平行注意力,利用对极线先验捕捉视差对齐对应关系,显著降低计算开销。在多个基准测试中,StereoWorld在立体一致性、视差准确性和相机运动保真度上优于强单目转立体流水线,生成速度提升3倍以上,视点一致性额外提高5%。此外,无需深度估计或修复即可实现端到端双目VR渲染,支持基于尺度化深度的具身策略学习,并兼容长视频蒸馏以实现长时间交互式立体合成。

原文摘要 · Abstract (English)

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry directly from disparity. To efficiently achieve consistent stereo generation, our approach introduces two key designs: (1) a unified camera-frame RoPE that augments latent tokens with camera-aware rotary positional encoding, enabling relative, view- and time-consistent conditioning while preserving pretrained video priors via a stable attention initialization; and (2) a stereo-aware attention decomposition that factors full 4D attention into 3D intra-view attention plus horizontal row attention, leveraging the epipolar prior to capture disparity-aligned correspondences with substantially lower compute. Across benchmarks, StereoWorld improves stereo consistency, disparity accuracy, and camera-motion fidelity over strong monocular-then-convert pipelines, achieving more than 3x faster generation with an additional 5% gain in viewpoint consistency. Beyond benchmarks, StereoWorld enables end-to-end binocular VR rendering without depth estimation or inpainting, enhances embodied policy learning through metric-scale depth grounding, and is compatible with long-video distillation for extended interactive stereo synthesis.

立体视频世界模型生成模型双目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。