提出新方法检测自动驾驶图像数据泄漏,提升模型评估可靠性。
Improving Image Data Leakage Detection in Automotive Software
- 基于计算实验构建泄漏检测机制,适用于汽车感知系统。
- 在Kitti数据集上首次发现未知的数据泄漏问题。
- 适合自动驾驶模型验证与工业级安全评估的研究者使用。
数据泄漏是机器学习/深度学习模型训练前划分训练集与测试集时常见的被忽视问题。泄漏会导致模型评估性能虚高,进而引发实时部署中的错误预测。然而,在需要图像输入的自动驾驶感知系统中,检测此类泄漏尤为困难。本研究基于工业合作伙伴沃尔沃汽车提供的Cirrus数据集开展计算实验,开发了一种数据泄漏检测方法,并在汽车领域广泛使用的公开基准数据集Kitti上进行了验证。结果表明,借助该方法成功识别出Kitti数据集中此前未被发现的数据泄漏现象。
原文摘要 · Abstract (English)
Data leakage is a very common problem that is often overlooked during splitting data into train and test sets before training any ML/DL model. The model performance gets artificially inflated with the presence of data leakage during the evaluation phase which often leads the model to erroneous prediction on real-time deployment. However, detecting the presence of such leakage is challenging, particularly in the object detection context of perception systems where the model needs to be supplied with image data for training. In this study, we conduct a computational experiment on the Cirrus dataset from our industrial partner Volvo Cars to develop a method for detecting data leakage. We then evaluate the method on another public dataset, Kitti, which is a popular and widely accepted benchmark dataset in the automotive domain. The results show that thanks to our proposed method we are able to detect data leakage in the Kitti dataset, which was previously unknown.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。