教小团队如何精准设计自动驾驶数据集,避免资源浪费。
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
- 先诊断问题是缺数据还是评测不完善,再选最小必要数据操作
- 只在更低成本方法无效时才采集新数据,节省资源
- 基于真实案例总结出可复用的数据集设计框架
高质量自动驾驶数据集推动了研究进展,但现有文献多描述数据内容,而非如何战略性地设计有影响力的数据集。这对资源有限的小型实验室和初创公司尤为不利。本文提出:构建高影响力数据集应始于诊断——判断研究问题是否受制于数据不足或评估缺陷,随后选择能填补缺口的最简数据操作,并仅在无更低成本方案时才采集新数据。我们以主要自动驾驶数据集的发展历程为案例,提炼出涵盖问题识别、操作选择、传感器配置与标注策略的战略框架,并通过我们的KITScenes数据集系列进行实证。数据集已开放获取:https://kitscenes.com/
原文摘要 · Abstract (English)
Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。