arXiv:2607.09402cs.LG2026-07

提出一套估算惯性传感器分类数据量的方法,让少样本也能高效训练。

Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

论文配图:Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification
图 1 · 摘自论文原文
  • 基于六组真实数据的实证分析,发现准确率随数据量呈对数增长。
  • 提出稳定点指标,证明模型可用更少数据达到实用稳定状态。
  • 适合需要规划数据采集的科研与工程人员参考使用。

深度学习模型对大规模惯性数据集的依赖,已成为人体活动识别和智能手机定位等任务中的主要瓶颈。这些领域需进行复杂、耗时且难以扩展的数据采集。目前尚无数据驱动的指导原则来确定达到目标精度所需的最小样本量。为此,本研究系统评估了惯性分类任务中学习曲线的收敛速率,提出统一框架,在二分类与多分类场景下分析分类性能,并推导出基于数据集规模的性能估算经验公式。在总计102.7小时的真实世界惯性数据上测试表明,无论任务复杂度如何,准确率均呈现一致的对数增长模式。基于此发现,我们定义了一个量化稳定点指标:即学习曲线在预设平均绝对百分比偏差内趋于稳定的样本量。分析显示,模型通常可在远低于传统经验法则建议的样本量下达到实际稳定。最终,我们提供一个可推广的框架,仅需小规模预实验即可外推总数据需求,优化采集投入与模型可靠性之间的权衡。该研究将范式从追求数据量转向数据效率,为惯性传感应用中的采集计划提供了切实可行、数据支持的指导。

原文摘要 · Abstract (English)

Deep learning models dependency on large-scale inertial datasets presents a significant bottleneck in inertial sensor-based classification tasks, such as human activity recognition and smartphone location recognition. In these domains, data collection requires massive recording campaigns that are complex, time-consuming, and difficult to scale. Currently, data-driven guidelines for determining the minimum sample size required to reach a desired accuracy level do not exist. To address this gap, this study presents a systematic empirical evaluation of learning curve convergence rates in inertial classification. We introduce a unified framework that analyzes classification performance under both binary and multi-class scenarios, and derive an empirical formula to estimate performance relative to dataset size. Testing across six diverse, real-world datasets totaling 102.7 hours of inertial measurements demonstrates that accuracy follows a consistent logarithmic growth pattern, regardless of task complexity. Leveraging this finding, we propose a quantitative stability point metric, defined as the sample size required for the learning curve to stabilize within a predefined mean absolute percentage deviation of its asymptotic maximum. Our analysis reveals that models often reach practical stability with substantially fewer samples than traditional heuristics suggest. Ultimately, we offer a generalizable framework to extrapolate total data requirements from small-scale pilot studies, optimizing the tradeoff between recording effort and model reliability. These findings shift the prevailing paradigm from maximizing data volume toward optimizing data efficiency, offering concrete, data-backed guidelines for planning recording campaigns in inertial sensing applications.

数据效率惯性传感学习曲线采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。