arXiv:2409.11307cs.CV2024-09中稿 · ICRA被引 1

让不同车的摄像头数据通用,提升3D视觉合成效果。

GS-Net: Heterogeneous Vehicle Data Reuse via Generalizable Plug-and-Play 3DGS Module

  • 用稀疏三维点云生成密集高斯分布,实现跨车辆数据复用。
  • 在插值和外推视角上分别提升2.08和1.86 dB PSNR。
  • 支持12个均匀分布摄像头,适配自动驾驶多传感器场景。

端到端自动驾驶日益依赖数据,但不同车辆间数据复用受限。每辆新车常需重新采集数据与训练,因摄像头位置、朝向和视场角差异所致。跨视角图像合成为跨平台数据复用提供可能,可通过现有传感器数据合成新配置下的图像。为此,我们提出GS-Net,一种轻量级即插即用模块,从稀疏SfM点云中聚合局部几何上下文,并在单次前向传播中将每个点扩展为多个稠密高斯原型,学习标准3DGS的跨场景泛化初始化,显著提升沿原始传感器轨迹的插值视图及新传感器位置的外推视图的渲染质量。现有方法与基准多集中于凸包内插值,而目标视角常位于训练相机凸包之外,该外推场景未被充分探索。为此,我们引入CARLA-NVS,首个专为跨传感器视角合成设计的基准。不同于通常仅含不超过八个固定摄像头的自动驾驶数据集,CARLA-NVS包含12个以30度方位角间隔均匀分布的摄像头,支持对插值与外推视角的受控评估。在CARLA-NVS上的实验表明,相较于标准3DGS,GS-Net在插值视图上提升2.08 dB PSNR,外推视图上提升1.86 dB,且初始化速度提升50倍。

原文摘要 · Abstract (English)

End-to-end autonomous driving is increasingly data-driven, yet data reuse across vehicles remains limited. Each new vehicle often requires additional data collection and retraining because camera translation, orientation, and field of view differ across sensor layouts. Cross-sensor view synthesis offers a promising route for cross-platform data reuse by synthesizing images under novel sensor configurations from existing sensor data. To realize this goal, we propose GS-Net, a lightweight plug-and-play module that aggregates local geometric context from sparse Structure-from-Motion (SfM) point clouds and expands each point into multiple dense Gaussian primitives in a single forward pass, learning a cross-scene generalizable initialization for standard 3DGS that improves rendering quality for both interpolated views along the original sensor trajectories and extrapolated views at new sensor positions. In such settings, target camera viewpoints often lie beyond the convex hull of training cameras, whereas existing methods and benchmarks predominantly evaluate in-hull interpolation, leaving the extrapolation regime underexplored. To enable quantitative evaluation under this setting, we introduce CARLA-NVS, the first benchmark explicitly designed for cross-sensor view synthesis. Unlike existing autonomous driving datasets that typically employ no more than eight cameras in fixed configurations, CARLA-NVS features 12 cameras uniformly distributed at 30-degree azimuth intervals, enabling controlled evaluation of both interpolated and extrapolated viewpoints. Experiments on CARLA-NVS show that GS-Net improves rendering quality by 2.08 dB PSNR on interpolated views and 1.86 dB on extrapolated views over standard 3DGS, while achieving a 50x faster initialization compared to MVS-based densification. These results offer a practical step toward scalable cross-vehicle data reuse in autonomous driving.

3DGS自动驾驶数据复用跨传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。