arXiv:2606.26087cs.CV2026-06

用多视角点追踪提升单目视频生成的几何与运动一致性

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

论文配图:MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
图 1 · 摘自论文原文
  • 通过多视角点追踪提供额外几何与运动监督信号
  • 在多个基准上实现最优几何一致性,相机轨迹精度领先
  • 适合需要精准运动与空间结构的4D视频生成研究者

从单目参考视频沿目标相机轨迹生成新视角视频,需同时保证几何一致性和运动保真度。现有基于显式3D表示的方法受限于现成重建模块对动态物体的几何误差;仅依赖相机条件的方法虽视觉质量高,但难以保持几何与运动一致性。本文提出MVTrack4Gen(多视角点追踪用于新视角生成),一种运动感知训练框架,利用多视角点追踪作为额外的几何与运动监督信号,作用于仅依赖相机条件的新视角视频扩散模型。关键发现:特定注意力层编码强对应关系,查询特征在跨视图和时间上关注几何对应位置,其错位导致运动不一致。据此,将这些特征引入辅助多视角追踪头,与扩散模型联合训练点追踪任务。通过显式强化此类运动感知对应,MVTrack4Gen使模型更准确跟随参考视图的运动并维持跨视图几何一致性。在多种基准上,本方法实现最先进的几何一致性,并达到具有竞争力的相机轨迹精度。

原文摘要 · Abstract (English)

Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos. In contrast, camera-conditioning-only methods can achieve high visual quality but often struggle to preserve geometric and motion consistency. In this work, we introduce MVTrack4Gen (Multi-View point Tracking for Novel-View Generation), a motion-aware training framework that leverages multi-view point tracking as an additional geometric and motion supervision signal for camera-conditioning-only novel-view video diffusion models. Our key finding is that specific attention layers encode strong correspondence cues, where query features attend to key features at geometrically corresponding locations across views and over time, and the misalignment of these correspondences causes motion inconsistency. Based on this observation, we route these features into an auxiliary multi-view tracking head and jointly train the diffusion model with a point-tracking objective. By explicitly strengthening these motion-aware correspondences, MVTrack4Gen improves existing models to better follow the motion in the reference view and maintain cross-view geometric consistency. Across diverse benchmarks, our method achieves state-of-the-art geometric consistency and competitive camera accuracy.

4D视频生成点追踪扩散模型几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。