arXiv:2512.08924cs.CV2025-12被引 33

用单一模型高效重建视频中动态场景的4D几何与运动

Efficiently Reconstructing Dynamic Scenes One D4RT at a Time

  • 统一变换器架构联合推断深度、时空对应关系和相机参数
  • 新查询机制避免密集解码,训练推理效率显著提升
  • 适合需要快速高精度4D重建的研究者与开发者

从视频中理解并重建动态场景的复杂几何与运动仍是计算机视觉中的重大挑战。本文提出D4RT,一种简单而强大的前馈模型,可高效完成此任务。D4RT采用统一的Transformer架构,从单个视频中联合推断深度、时空对应关系和完整相机参数。其核心创新在于一种新型查询机制,绕过了密集帧解码的高昂计算开销以及多任务专用解码器的复杂管理。该解码接口使模型能独立且灵活地探测任意时空位置的3D坐标。结果表明,该方法轻量且高度可扩展,实现了显著高效的训练与推理。我们在广泛的4D重建任务中验证了其性能,达到新的最优水平。

原文摘要 · Abstract (English)

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently solve this task. D4RT utilizes a unified transformer architecture to jointly infer depth, spatio-temporal correspondence, and full camera parameters from a single video. Its core innovation is a novel querying mechanism that sidesteps the heavy computation of dense, per-frame decoding and the complexity of managing multiple, task-specific decoders. Our decoding interface allows the model to independently and flexibly probe the 3D position of any point in space and time. The result is a lightweight and highly scalable method that enables remarkably efficient training and inference. We demonstrate that our approach sets a new state of the art, outperforming previous methods across a wide spectrum of 4D reconstruction tasks. We refer to the project webpage for animated results: https://d4rt-paper.github.io/.

4D重建Transformer视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。