用AI+人工协作加速自动驾驶多传感器数据标注
Semi-Automated Data Annotation in Multisensor Datasets for Autonomous Vehicle Testing
- AI先生成初始标注,人工修正并反馈重训模型
- 节省大量时间,保证跨传感器标注一致性
- 适合自动驾驶数据集构建与研究团队使用
本报告介绍了在DARTS项目中设计并实现的半自动化数据标注流程,旨在创建大规模、多模态的波兰驾驶场景数据集。手动标注此类异构数据成本高且耗时。为应对这一挑战,所提方案采用人机协同模式,结合人工智能与人类专业知识,降低标注成本和耗时。系统自动产生初始标注,支持模型迭代重训练,并集成数据匿名化与领域自适应技术。核心依赖3D目标检测算法生成初步标注。整体工具与方法显著节省时间,确保不同传感器模态间标注的一致性与高质量。该方案直接支持DARTS项目,加快了标准化格式大尺寸标注数据集的准备,强化了波兰自动驾驶研究的技术基础。
原文摘要 · Abstract (English)
This report presents the design and implementation of a semi-automated data annotation pipeline developed within the DARTS project, whose goal is to create a large-scale, multimodal dataset of driving scenarios recorded in Polish conditions. Manual annotation of such heterogeneous data is both costly and time-consuming. To address this challenge, the proposed solution adopts a human-in-the-loop approach that combines artificial intelligence with human expertise to reduce annotation cost and duration. The system automatically generates initial annotations, enables iterative model retraining, and incorporates data anonymization and domain adaptation techniques. At its core, the tool relies on 3D object detection algorithms to produce preliminary annotations. Overall, the developed tools and methodology result in substantial time savings while ensuring consistent, high-quality annotations across different sensor modalities. The solution directly supports the DARTS project by accelerating the preparation of large annotated dataset in the project's standardized format, strengthening the technological base for autonomous vehicle research in Poland.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。