arXiv:2509.12683cs.CV2025-09被引 6

构建高保真立体数据集,提升自动驾驶模型泛化能力

StereoCarla: A High-Fidelity Driving Dataset for Generalizable Stereo

  • 基于CARLA仿真器生成多视角、多环境的立体图像数据
  • 在4个基准数据集上训练的模型精度超越11个现有数据集
  • 适合研究自动驾驶立体匹配与泛化性能的学者使用

立体匹配对自动驾驶和机器人深度感知至关重要。尽管近年学习型算法和合成数据集推动了该领域发展,但模型泛化能力仍受限于训练数据多样性不足。为此,我们提出StereoCarla——一个专为自动驾驶设计的高保真合成立体数据集。基于CARLA模拟器,该数据集涵盖多种相机配置(如不同基线、视角、传感器位置)及多变环境条件(光照、天气、道路几何)。我们在四个标准评估数据集(KITTI2012、KITTI2015、Middlebury、ETH3D)上进行跨域实验,结果表明:在StereoCarla上训练的模型,在多个基准测试中均优于在11个现有数据集上训练的模型。此外,将其融入多数据集训练可显著提升泛化精度,体现其兼容性与可扩展性。该数据集为真实、多样、可控环境下立体算法的研发与评估提供了重要基准,助力更鲁棒的自动驾驶深度感知系统。代码与数据将公开于https://github.com/XiandaGuo/OpenStereo 和 https://xiandaguo.net/StereoCarla。

原文摘要 · Abstract (English)

Stereo matching plays a crucial role in enabling depth perception for autonomous driving and robotics. While recent years have witnessed remarkable progress in stereo matching algorithms, largely driven by learning-based methods and synthetic datasets, the generalization performance of these models remains constrained by the limited diversity of existing training data. To address these challenges, we present StereoCarla, a high-fidelity synthetic stereo dataset specifically designed for autonomous driving scenarios. Built on the CARLA simulator, StereoCarla incorporates a wide range of camera configurations, including diverse baselines, viewpoints, and sensor placements as well as varied environmental conditions such as lighting changes, weather effects, and road geometries. We conduct comprehensive cross-domain experiments across four standard evaluation datasets (KITTI2012, KITTI2015, Middlebury, ETH3D) and demonstrate that models trained on StereoCarla outperform those trained on 11 existing stereo datasets in terms of generalization accuracy across multiple benchmarks. Furthermore, when integrated into multi-dataset training, StereoCarla contributes substantial improvements to generalization accuracy, highlighting its compatibility and scalability. This dataset provides a valuable benchmark for developing and evaluating stereo algorithms under realistic, diverse, and controllable settings, facilitating more robust depth perception systems for autonomous vehicles. Code can be available at https://github.com/XiandaGuo/OpenStereo, and data can be available at https://xiandaguo.net/StereoCarla.

立体匹配自动驾驶合成数据泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。