arXiv:2606.19712cs.LGcs.CV2026-06

基于数据特性快速选出适合少类别任务的轻量模型

Efficient Neural Network Model Selection for Few-Class Application Datasets

论文配图:Efficient Neural Network Model Selection for Few-Class Application Datasets
图 1 · 摘自论文原文
  • 提出用数据侧属性衡量分类难度,指导模型选择
  • 比传统方法快6到29倍,模型可缩小42%且性能不降
  • 特别适合移动机器人、无人机等资源受限场景

尽管大量研究致力于开发高性能神经网络,但对数据集特性如何指导高效模型选择的关注较少。神经模型通常在含数千类的数据集上评估,而许多实际应用仅涉及少于十个类别。针对这一常见但研究不足的场景,我们提出一种基于数据侧特性的分类难度度量,证明其可显著提升少类别数据集上的模型选择效率,该现象称为‘少类别独特性’。该度量使模型与数据集比较速度比重复训练测试快6至29倍。基于此,我们将模型家族规模扩展至小于已有最小模型,实现相似精度下更高效率,例如在移动机器人任务中模型比YOLOv5-nano小42%。面向资源受限的应用,我们在移动机器人、无人机和物联网场景中验证了少类别模型选择的实际收益,无需牺牲性能即可大幅提升效率。

原文摘要 · Abstract (English)

While much effort has focused on developing and benchmarking high-performance neural networks, less attention has been given to how dataset properties, known to practitioners, can guide efficient model selection. Neural models are typically evaluated on datasets with thousands of classes, yet many real-world applications involve fewer than ten. To address this understudied but common setting, we develop a measure of classification difficulty based on data-side properties and show how it enables more efficient model selection for few-class datasets, where traditional approaches are less effective. We term this phenomenon "few-class distinctiveness". Our metric allows comparison of models and datasets 6 to 29$\times$ faster than repeated training and testing. Leveraging this insight, we extend scaled model families below the smallest published models, achieving greater efficiency at similar accuracy, for example models up to 42% smaller than YOLOv5-nano for a mobile robot task. Targeting resource-constrained applications, we demonstrate few-class model selection across mobile robot, drone, and IoT scenarios, highlighting practical gains in efficiency without sacrificing performance.

模型选择少类别轻量化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。