混合真实与合成数据可提升自动驾驶目标检测的鲁棒性。
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
- 用真实与合成数据混合训练,提升模型泛化能力。
- 合成数据虽有分布差异,但组合使用能增强检测性能。
- 适用于自动驾驶视觉与激光雷达系统的模型优化。
自动驾驶系统日益普及,对大规模高质量数据集的需求随之增长。合成数据因其成本低、标注精准且可模拟极端场景,成为补充真实数据的可行方案。然而,合成数据可能引入分布偏差,影响模型在真实环境中的表现。为评估其价值与局限,我们基于多组真实数据集及比特科技公司生成的合成数据,开展了受控实验,覆盖摄像头与激光雷达两种传感器模态,研究2D与3D目标检测任务。对比了仅用真实数据、仅用合成数据、以及混合数据训练的模型在鲁棒性与泛化能力上的表现。结果表明,融合真实与合成数据能显著提升模型性能,验证了合成数据在推动自动驾驶技术发展中的潜力。
原文摘要 · Abstract (English)
The increasing applications of autonomous driving systems necessitates large-scale, high-quality datasets to ensure robust performance across diverse scenarios. Synthetic data has emerged as a viable solution to augment real-world datasets due to its cost-effectiveness, availability of precise ground-truth labels, and the ability to model specific edge cases. However, synthetic data may introduce distributional differences and biases that could impact model performance in real-world settings. To evaluate the utility and limitations of synthetic data, we conducted controlled experiments using multiple real-world datasets and a synthetic dataset generated by BIT Technology Solutions GmbH. Our study spans two sensor modalities, camera and LiDAR, and investigates both 2D and 3D object detection tasks. We compare models trained on real, synthetic, and mixed datasets, analyzing their robustness and generalization capabilities. Our findings demonstrate that the use of a combination of real and synthetic data improves the robustness and generalization of object detection models, underscoring the potential of synthetic data in advancing autonomous driving technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。