用少量移动相机实现低成本高精度4D动态重建
4D Reconstruction from Sparse Dynamic Cameras

- 通过跨相机特征匹配+相机内点追踪,保证时空一致性
- 在真实场景数据集上,动态区域重建质量显著优于基线
- 适合体育、演唱会等低成本多机位视频制作
尽管单目动态相机的4D重建近年取得进展,仍受限于深度模糊。本文提出一种稀疏动态相机方案:少量独立运动的相机拍摄同一主体,兼顾低成本与多视角约束,适用于体育、演唱会、电视节目等真实场景。实验表明,直接套用现有单目或密集固定相机方法无法解决跨视图和时间的复杂不一致性。为此,我们提出一种简单有效的3D轨迹初始化方法,通过整合跨相机特征匹配与相机内点跟踪来确保时空一致性;同时引入噪声鲁棒的深度排序正则化损失和时空多样性批次采样策略,提升优化稳定性和跨视图泛化能力。此外,为填补该任务缺乏标准基准的空白,我们构建了LetCamsGo数据集,包含4种环境下的5个真实视频序列,由三台独立移动相机与一台固定相机同步拍摄。在LetCamsGo上的全面评测表明,所提框架在动态区域的4D重建质量显著优于基线,为野外低成本4D重建开辟新路径。
原文摘要 · Abstract (English)
Although dynamic 3D (i.e., 4D) reconstruction from a monocular dynamic camera has recently advanced, it remains fundamentally limited by depth ambiguity. In this paper, we focus on an alternative practical way, i.e., sparse dynamic camera setup, where a handful of independently moving cameras capture the same subjects. While keeping capture costs low, this setup introduces multi-view constraints and remains practical for real-world video production such as sports, concerts, and TV shows. Despite its potential, our experiments show that naive extensions of existing monocular or dense-fixed camera-based methods are insufficient since they fail to resolve the complex spatiotemporal inconsistencies across views and time. To fill this gap, we propose a simple yet effective 3D track initialization method designed to ensure spatiotemporal consistency by integrating inter-camera feature matching with intra-camera point tracking. Additionally, we incorporate a noise-robust depth-ordering regularization loss and a spatiotemporally diverse batch sampling strategy to enhance optimization stability and cross-view generalization. Furthermore, to address the lack of standardized benchmarks for this task, we introduce LetCamsGo, a new real-world video dataset with 5 sequences across 4 diverse environments, recorded by three independently moving cameras and one fixed camera. Comprehensive benchmarking on LetCamsGo demonstrated that our proposed framework improves 4D reconstruction quality in dynamic regions compared with baselines, paving the way for a low-cost 4D reconstruction paradigm in the wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。