构建多车多视角驾驶数据集,推动动态城市场景新视图合成评测
MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

- 同步采集汽车、电动自行车、无人机三类载体的多视角视频
- 12000张图像覆盖50个高质量场景,支持大视角变化下的合成测试
- 揭示当前方法在大视角差异下性能下降,适合自动驾驶视觉研究者
可微渲染推动了新视图合成(NVS)的发展,但在真实驾驶场景中仍面临视角稀疏、动态物体和多轨迹数据有限等挑战。本文提出多视角多车辆(MV2)数据集及基准,用于评估动态城市场景中大视角变化下的NVS模型表现。MV2包含汽车、电动自行车和无人机同步采集的数据,各载体沿不同但同步的轨迹行驶。以一辆车的摄像头序列训练,另一辆车的序列测试,实现比现有单轨迹数据集更大的视角变化评估。所有序列通过结构光重建注册,相机位姿经人工像素级对应标注验证,共生成50个高质量场景,含12000张图像。对近期NVS与相机位姿估计方法的基准测试表明,随着视角差异增大,NVS性能显著下降,且前馈式位姿估计算法明显落后于优化方法,凸显了MV2作为驾驶场景下NVS严格测试平台的价值。数据集、基准协议与项目资源详见https://mv2-dataset.github.io/。
原文摘要 · Abstract (English)
Differentiable rendering has advanced novel view synthesis (NVS), yet applying it to real-world driving remains difficult due to sparse capture viewpoints, dynamic objects, and limited multi-trajectory data. We introduce the Multi-View Multi-Vehicle (MV2) dataset and benchmark for evaluating NVS models under large viewpoint changes in dynamic urban scenes. MV2 features synchronized captures from a car, scooter, and drone, each following distinct yet synchronized trajectories. Training NVS methods on one vehicle's camera stream and testing on another enables evaluation under substantially larger viewpoint variations than existing single-trajectory datasets. All sequences are registered via Structure-from-Motion and camera poses verified using manual pixel-level correspondence annotations, yielding 50 high-quality scenes with 12000 images. Benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed-forward pose estimators notably lag behind optimization-based approaches, highlighting MV2 as a rigorous testbed for NVS in driving. The dataset, benchmark protocol, and project resources are available at https://mv2-dataset.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。