单个查询统一重建动态4D场景,支持任意视图对与稀疏推理。
UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

- 用查询条件控制多帧编码,解码时动态选源目标视图
- 在WorldTrack上动态点重建与光流估计均达最佳宏观平均性能
- 无需固定时长嵌入,支持稀疏查询与批量密集重建
动态4D场景重建需联合估计对应关系、几何结构、物体运动和相机运动。现有前馈方法通常预测稠密任务特定图或独立处理源-目标帧对,导致稀疏查询时计算冗余,且不同帧对间特征复用有限。我们提出UniQuery4R,一种查询条件化的框架:将多帧片段一次性编码,仅在解码时通过源-目标交叉注意力选择源视图、目标视图及连续源图像坐标。每个查询联合预测目标对应关系、目标时刻3D位置、场景流,以及源深度,同时每视图独立估计相机参数。该设计使编码片段可跨任意源-目标选择复用,支持稀疏推理与批量密集重建,且无需依赖固定片段长度的时序嵌入。此外,我们引入方向-幅度参数化场景流,并对移动点与静态点分别施加监督。在评估方法中,UniQuery4R在WorldTrack数据集上动态点重建与场景流估计的宏平均表现最优。
原文摘要 · Abstract (English)
Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed-forward methods typically predict dense task-specific maps or independently process source-target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention. Each query jointly predicts target correspondence, target-time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source-target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction-magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro-average results on WorldTrack for both scene-flow estimation and dynamic-point reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。