开源工具助力智能交通系统从海量数据中筛选罕见目标
Mcity Data Engine: Iterative Model Improvement Through Open-Vocabulary Data Selection
- 基于开放词汇表迭代筛选稀有和新类样本
- 支持从采集到部署的完整数据开发流程
- 适合自动驾驶与感知研究者使用
随着数据量持续增长,为机器学习模型选择并标注合适样本变得愈发困难,尤其在大量无标签数据中识别长尾类别更为挑战。这一问题在智能交通系统(ITS)中尤为突出,因车辆车队与路边感知系统产生海量原始数据。尽管工业界已有专有数据引擎用于迭代式数据选择与模型训练,但研究者和开源社区缺乏公开可用的系统。本文提出Mcity Data Engine,提供从数据采集到模型部署的全流程模块化支持,聚焦通过开放词汇表机制识别罕见及新型类别。所有代码已开源,采用MIT许可证,可在GitHub获取:https://github.com/mcity/mcity_data_engine
原文摘要 · Abstract (English)
With an ever-increasing availability of data, it has become more and more challenging to select and label appropriate samples for the training of machine learning models. It is especially difficult to detect long-tail classes of interest in large amounts of unlabeled data. This holds especially true for Intelligent Transportation Systems (ITS), where vehicle fleets and roadside perception systems generate an abundance of raw data. While industrial, proprietary data engines for such iterative data selection and model training processes exist, researchers and the open-source community suffer from a lack of an openly available system. We present the Mcity Data Engine, which provides modules for the complete data-based development cycle, beginning at the data acquisition phase and ending at the model deployment stage. The Mcity Data Engine focuses on rare and novel classes through an open-vocabulary data selection process. All code is publicly available on GitHub under an MIT license: https://github.com/mcity/mcity_data_engine
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。