用GitOps管理铁路自动驾驶数据标注,实现全流程可追溯
A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations

- 以Data-as-Code理念构建轻量级标注元数据管理架构
- 通过CI/CD与静态站点生成实现自动化的数据集概览
- 适合自动驾驶系统研发团队和合规性要求高的项目
自动化列车运行(GoA3-GoA4)依赖于鲁棒的基于AI的感知系统,能在真实环境下可靠检测障碍物和铁路特定物体。这类现代AI方法的有效性高度依赖大规模、高质量且动态更新的标注数据集。然而,管理元数据、维护溯源记录并追踪标注的迭代演化带来了巨大的基础设施和监管压力。现有集中式数据目录常面临操作开销大、与开发流程集成差、文档滞后等问题。本文提出一种创新的轻量级GitOps驱动型元数据管理架构。通过采用Data-as-Code原则、CI/CD流水线和静态站点生成(SSG),该方法建立了一条无缝的开发者导向工作流,确保可追溯性,强制执行严格监管合规,并自动生成高性能的数据集概览。
原文摘要 · Abstract (English)
Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based perception systems capable of reliably detecting obstacles and railway-specific objects under real-world conditions. The effectiveness of these modern artificial intelligence approaches depends heavily on large-scale, high-quality, and highly dynamic annotated datasets. However, managing metadata, maintaining provenance, and tracking the iterative evolution of these annotations impose significant infrastructural and regulatory requirements. Existing monolithic data catalogs often suffer from massive operational overhead, poor integration into developer workflows, and severe documentation drift. This paper introduces an innovative, lightweight GitOps-based architecture for metadata management. By leveraging Data-as-Code principles, Continuous Integration/Continuous Deployment (CI/CD) pipelines, and Static Site Generation (SSG), the proposed approach establishes a seamless, developer-centric workflow. This ensures an traceability, enforces strict regulatory compliance, and automatically generates a highly performant dataset overview.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。