arXiv:2606.15534cs.CV2026-06

用3D点轨迹实现视角切换时视频内容稳定一致的生成

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

论文配图:Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
图 1 · 摘自论文原文
  • 通过配对3D点轨迹建立源与目标视图的显式时空对应关系
  • 在400个视频上实现最高视觉质量,旋转误差降低30%-65%
  • 适合需要精确相机控制和动态场景重渲染的研究者

从新视角重渲染现有视频需遵循指定相机轨迹,同时保持各帧中场景外观与动态的一致性。现有方法依赖逐帧姿态嵌入、噪声点云渲染或隐式学习对应关系,均未能提供源与目标像素间显式、连续的时空连接。我们提出Track2View,将视频扩散变换器基于配对3D点轨迹进行条件控制:即场景点在源视图与目标视图中的稀疏轨迹投影。这些轨迹天然具有时间连续性,明确编码了内容应出现在何处及何时。Track2View的核心是双视图轨迹调节器,通过无参数几何操作与可学习时间聚合,将视觉上下文从源视图传递至目标视图,无需记忆特定运动即可泛化至任意相机轨迹。我们还引入数据整理流程,通过在时序拼接的多相机视图对上运行3D点追踪器,提取一一对应的轨迹。在涵盖静态与动态场景的400视频基准测试中,Track2View在视觉质量、视角同步性和相机精度方面均达到最先进水平,相比领先基线,旋转误差降低30%-65%,平移误差降低61%-72%。

原文摘要 · Abstract (English)

Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the original scene across every frame. Existing methods rely on per-frame pose embeddings, noisy point-cloud renderings, or implicit learned correspondences, none of which provides an explicit, temporally continuous link between source and target pixels. We propose Track2View, which conditions a video diffusion transformer on paired 3D point tracks: sparse trajectories of scene points projected into both the source and target camera views. These tracks provide explicit spatiotemporal correspondences that are temporally continuous by construction, encoding what content should appear where and when. At the core of Track2View is a dual-view track conditioner that transfers visual context from source to target view through parameter-free geometric operations and learned temporal aggregation, ensuring generalization to arbitrary camera trajectories without memorizing specific motions. We further introduce a data curation pipeline that extracts one-to-one track correspondences by running a 3D point tracker on temporally concatenated multi-camera view pairs. On a 400-video benchmark spanning static and dynamic scenes, Track2View achieves state-of-the-art results across visual quality, view synchronization, and camera accuracy, reducing rotation error by 30-65% and translation error by 61-72% relative to leading baselines. Project page is available at this https URL: https://qjizhi.github.io/track2view

视频生成4D一致性相机控制点轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。