分析423个类脑数据集,揭示其规模混乱、标准缺失与合成数据隐患。
LAND: A Longitudinal Analysis of Neuromorphic Datasets
- 梳理423个类脑数据集,剖析任务结构与数据特性
- 发现数据规模持续增长,但访问难、标准化差
- 提出元数据集概念,缓解数据冗余与偏差问题
类脑工程面临数据困境。尽管过去十年中类脑数据集数量激增,但大量研究仍宣称需要更多、更大的数据集。这不仅源于现代深度学习对海量数据的需求,更因现有数据集存在难以获取、用途不清、任务不明等问题,且实际下载与使用存在困难。本文首次系统梳理超过423个类脑数据集,分析其任务本质与数据结构,揭示其在规模、标准化及可访问性方面的挑战。同时指出单个数据集规模持续扩大,处理复杂度提升。更值得关注的是合成数据(仿真或视频转事件)的兴起,虽利于算法验证,却可能误导新应用探索。本文提出元数据集(meta-datasets)概念,通过整合已有数据减少新增数据需求,并降低因数据与任务定义不当带来的偏差风险。
原文摘要 · Abstract (English)
Neuromorphic engineering has a data problem. Despite the meteoric rise in the number of neuromorphic datasets published over the past ten years, the conclusion of a significant portion of neuromorphic research papers still states that there is a need for yet more data and even larger datasets. Whilst this need is driven in part by the sheer volume of data required by modern deep learning approaches, it is also fuelled by the current state of the available neuromorphic datasets and the difficulties in finding them, understanding their purpose, and determining the nature of their underlying task. This is further compounded by practical difficulties in downloading and using these datasets. This review starts by capturing a snapshot of the existing neuromorphic datasets, covering over 423 datasets, and then explores the nature of their tasks and the underlying structure of the presented data. Analysing these datasets shows the difficulties arising from their size, the lack of standardisation, and difficulties in accessing the actual data. This paper also highlights the growth in the size of individual datasets and the complexities involved in working with the data. However, a more important concern is the rise of synthetic datasets, created by either simulation or video-to-events methods. This review explores the benefits of simulated data for testing existing algorithms and applications, highlighting the potential pitfalls for exploring new applications of neuromorphic technologies. This review also introduces the concepts of meta-datasets, created from existing datasets, as a way of both reducing the need for more data, and to remove potential bias arising from defining both the dataset and the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。