arXiv:2601.00678cs.CV2026-01被引 1

单图生成可控视频,用3D高斯模型实现相机轨迹驱动的动态画面

Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians

  • 用3D高斯表示场景,单次前向传播生成带运动的视频
  • 在KITTI等4个数据集上达到最优画质与推理速度
  • 适合需要精准相机控制的自动驾驶、虚拟拍摄场景

人类仅凭一张图像就能预测场景未来的动态。能够模拟这一能力的视频生成模型是智能系统的关键组成部分。现有方法虽提升了单图生成视频的时间连贯性和3D一致性,但普遍缺乏对相机路径等用户操作的鲁棒控制,限制了实际应用。多数相机可控生成模型难以准确建模相机运动,保持时间一致性并维持几何完整性。利用显式中间3D表示可解决此问题,实现与给定相机轨迹对齐的连贯视频生成。尽管已有方法采用3D点云渲染场景并在后续阶段引入物体运动,但两步流程仍难保证完全的时间一致性,即便能精确控制相机移动。本文提出新框架:在单次前向传播中构建3D高斯场景表示,并采样合理的物体运动,实现无需迭代去噪的快速相机引导视频生成。在KITTI、Waymo、RealEstate10K和DL3DV-10K数据集上的大量实验表明,该方法在视频质量与推理效率方面均达到当前最佳。项目主页见 https://melonienimasha.github.io/Pixel-to-4D-Website。

原文摘要 · Abstract (English)

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence and 3D consistency in single-image-conditioned video generation. However, these methods often lack robust user controllability, such as modifying the camera path, limiting their applicability in real-world applications. Most existing camera-controlled image-to-video models struggle with accurately modeling camera motion, maintaining temporal consistency, and preserving geometric integrity. Leveraging explicit intermediate 3D representations offers a promising solution by enabling coherent video generation aligned with a given camera trajectory. Although these methods often use 3D point clouds to render scenes and introduce object motion in a later stage, this two-step process still falls short in achieving full temporal consistency, despite allowing precise control over camera movement. We propose a novel framework that constructs a 3D Gaussian scene representation and samples plausible object motion, given a single image in a single forward pass. This enables fast, camera-guided video generation without the need for iterative denoising to inject object motion into render frames. Extensive experiments on the KITTI, Waymo, RealEstate10K and DL3DV-10K datasets demonstrate that our method achieves state-of-the-art video quality and inference efficiency. The project page is available at https://melonienimasha.github.io/Pixel-to-4D-Website.

图像转视频3D高斯相机控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。