构建真实与合成数据集,提升动态相机参数估计精度。
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics

- 生成441K+帧的合成视频,含每帧精确内参标注。
- 扩展真实数据至514K+帧,覆盖更广场景与运动类型。
- 合成数据可有效提升现有方法在真实视频中的表现。
相机内参对从2D视频恢复3D结构至关重要。然而,多数3D算法假设内参固定,这在真实视频中常不成立。因此,从RGB图像估计逐帧内参对提升3D方法鲁棒性至关重要。InFlux首次建立了带有逐帧真实内参标注的真实世界基准。但现有方法仍不准确,主要受限于:(i) 训练数据稀缺且内参多样性不足;(ii) 基准数据场景与相机运动多样性有限,难以全面评估方法。为解决上述问题,我们提出InFlux++,包含两个部分。InFlux++ Synth是大规模程序化生成的合成视频数据集,包含441,000+标注帧,来自1841段高分辨率视频,提供逐帧精确内参标注;部分子集还包含逐帧位姿、深度和法向量。视频通过变焦、对焦变化以及真实渲染效果(如镜头畸变、散焦模糊)实现丰富的内参多样性。InFlux++ Real是真实世界基准的扩展,新增514,000+帧,覆盖334段高分辨率视频,涵盖更广泛的场景和相机运动。在InFlux++ Synth上微调现有内参预测方法,显著提升了在InFlux++ Real和InFlux上的焦距估计性能,表明基于合成数据的监督对基于RGB的内参预测具有前景。数据集、基准、代码、视频、提交说明及实时排行榜详见 https://influx.cs.princeton.edu/。
原文摘要 · Abstract (English)
Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assumption that often fails for real-world in-the-wild videos. Consequently, estimating per-frame intrinsics from RGB images is critical for making 3D methods robust to videos with dynamic intrinsics. InFlux previously advanced this research direction by establishing the first real-world benchmark with per-frame ground truth intrinsics for dynamic intrinsics videos. Nevertheless, existing methods remain inaccurate due to two obstacles: (i) training data is scarce and lacks intrinsics diversity; and (ii) benchmarks, including InFlux, have limited scene and camera motion diversity, making it difficult to properly evaluate methods. To address both gaps, we present InFlux++, consisting of two components. InFlux++ Synth is a large-scale procedurally generated synthetic video dataset with 441K+ annotated frames from 1841 high-resolution videos, providing accurate per-frame ground truth intrinsics for training dynamic intrinsics prediction models; a subset also includes per-frame pose, depth, and normals. The videos feature rich intrinsics diversity through changes in camera zoom and focus, as well as dynamic objects and realistic rendering effects such as lens distortion and defocus blur. InFlux++ Real is a large-scale real-world benchmark that extends InFlux with 514K+ newly captured frames across 334 high-resolution videos, spanning a wider range of scenes and camera motions. Finetuning existing intrinsics prediction methods on InFlux++ Synth consistently improves focal length estimation across both InFlux++ Real and InFlux, suggesting that synthetic supervision is promising for RGB-based intrinsics prediction. For the dataset, benchmark, code, videos, submission instructions, and live leaderboard, please visit https://influx.cs.princeton.edu/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。