车企感知系统中数据泄露隐患被忽视,实操者缺乏统一应对策略。
Data Leakage in Automotive Perception: Practitioners' Insights
- 分角色认知差异:算法与验证人员对泄露理解不同
- 检测靠经验异常,无专用工具,预防依赖口头交流
- 需跨角色协作建立统一定义与可追溯数据流程
数据泄露是指训练与评估数据集之间无意的信息传递,在自动驾驶感知等安全关键系统中会严重威胁机器学习模型的可靠性。尽管学术界已广泛关注此问题,但工业界实践者如何认知和应对仍不明确。本研究通过与10名负责汽车感知功能设计、开发和验证的工程师进行半结构化访谈,采用反思性主题分析方法发现:数据泄露的认知在工程团队中普遍存在但分散于岗位边界——机器学习工程师将其视为数据划分或验证问题,而设计与验证角色则从场景覆盖性和代表性角度理解。检测通常基于通用经验判断和性能异常,而非专用工具。预防措施多依赖经验积累和知识共享。结果表明,泄露控制本质上是跨角色、跨流程的协同难题。研究呼吁加强机器学习可靠性工程,推动建立共享定义、可追溯的数据实践以及持续的跨角色沟通,以在汽车机器学习开发中制度化数据泄露意识。
原文摘要 · Abstract (English)
Data leakage is the inadvertent transfer of information between training and evaluation datasets that poses a subtle, yet critical, risk to the reliability of machine learning (ML) models in safety-critical systems such as automotive perception. While leakage is widely recognized in research, little is known about how industrial practitioners actually perceive and manage it in practice. This study investigates practitioners' knowledge, experiences, and mitigation strategies around data leakage through ten semi-structured interviews with system design, development, and verification engineers working on automotive perception functions development. Using reflexive thematic analysis, we identify that knowledge of data leakage is widespread and fragmented along role boundaries: ML engineers conceptualize it as a data-splitting or validation issue, whereas design and verification roles interpret it in terms of representativeness and scenario coverage. Detection commonly arises through generic considerations and observed performance anomalies rather than implying specific tools. However, data leakage prevention is more commonly practiced, which depends mostly on experience and knowledge sharing. These findings suggest that leakage control is a socio-technical coordination problem distributed across roles and workflows. We discuss implications for ML reliability engineering, highlighting the need for shared definitions, traceable data practices, and continuous cross-role communication to institutionalize data leakage awareness within automotive ML development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。