4D-VGGT用分治法统一建模动态场景的时空几何,提升泛化能力。
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
- 分治式时空建模:空间与时间特征分别处理,避免异构融合失配
- 支持任意视角和时序输入,适配多场景动态几何估计任务
- 多任务头设计,可同时完成多种几何预测,适合通用视觉系统
我们研究动态场景几何估计这一挑战性任务,需同时建模空间与时间特征。现有方法常将二者对齐至统一潜在空间,但因空间与时间特征本质差异,易导致表征失配。为此,我们提出4D-VGGT——一种具备分治式时空感知的通用基础模型。模型包含三方面:1)多设置输入:设计自适应视觉网格,支持任意数量视角与时间步的输入序列;2)多层级表征:采用跨视图全局融合实现空间表示,跨时间局部融合实现时间表示;3)多任务预测:在时空表征后接多个任务专用头,实现动态场景的全面几何估计。该统一框架增强了特征区分度与应用普适性。我们整合多个几何数据集进行训练,并在多个动态场景几何基准上开展大量实验,验证了方法在多种任务中的有效性。
原文摘要 · Abstract (English)
We investigate a challenging task of dynamic scene geometry estimation, which requires representing both spatial and temporal features. Typically, existing methods align the two features into a unified latent space to model scene geometry. However, this unified paradigm suffers from potential mismatched representation due to the heterogeneous nature between spatial and temporal features. In this work, we propose 4D-VGGT, a general foundation model with divide-and-conquer spatiotemporal representation for dynamic scene geometry. Our model is divided into three aspects: 1) Multi-setting input. We design an adaptive visual grid that supports input sequences with arbitrary numbers of views and time steps. 2) Multi-level representation. We propose a cross-view global fusion for spatial representation and a cross-time local fusion for temporal representation. 3) Multi-task prediction. We append multiple task-specific heads to spatiotemporal representations, enabling a comprehensive visual geometry estimation for dynamic scenes. Under this unified framework, these components enhance the feature discriminability and application universality of our model for dynamic scenes. In addition, we integrate multiple geometry datasets to train our model and conduct extensive experiments to verify the effectiveness of our method across various tasks on multiple dynamic scene geometry benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。