arXiv:2506.07612cs.CV2025-06被引 4

用虚拟传感器数据提升动作识别,尤其在数据少时效果显著。

Scaling Human Activity Recognition: A Comparative Evaluation of Synthetic Data Generation and Augmentation Techniques

  • 通过视频和语言生成虚拟IMU数据,替代真实采集
  • 在小数据场景下,虚拟数据比真实或增强数据性能更好
  • 为不同需求提供生成策略选择建议,适合数据受限研究

人体动作识别(HAR)常受限于标注数据稀缺,因真实数据收集成本高、难度大。为缓解此问题,近年研究探索通过跨模态迁移生成虚拟惯性测量单元(IMU)数据。视频与语言驱动的生成管道各有优势,但假设条件与计算开销不同,且其相对于传统传感器级数据增强的效果尚不明确。本文首次直接对比两种虚拟IMU生成方法与经典数据增强技术。我们构建了一个大规模虚拟IMU数据集,涵盖来自Kinetics-400的100种多样化动作,并在22个身体部位模拟传感器信号。三种生成策略在四个基准HAR数据集(UTD-MHAD, PAMAP2, HAD-AW)上,使用四种主流模型进行评估。结果表明,虚拟IMU数据在有限数据条件下显著优于仅使用真实或增强数据,大幅提升识别性能。本文提供实际策略选择指导,并揭示各方法的优劣势。

原文摘要 · Abstract (English)

Human activity recognition (HAR) is often limited by the scarcity of labeled datasets due to the high cost and complexity of real-world data collection. To mitigate this, recent work has explored generating virtual inertial measurement unit (IMU) data via cross-modality transfer. While video-based and language-based pipelines have each shown promise, they differ in assumptions and computational cost. Moreover, their effectiveness relative to traditional sensor-level data augmentation remains unclear. In this paper, we present a direct comparison between these two virtual IMU generation approaches against classical data augmentation techniques. We construct a large-scale virtual IMU dataset spanning 100 diverse activities from Kinetics-400 and simulate sensor signals at 22 body locations. The three data generation strategies are evaluated on benchmark HAR datasets (UTD-MHAD, PAMAP2, HAD-AW) using four popular models. Results show that virtual IMU data significantly improves performance over real or augmented data alone, particularly under limited-data conditions. We offer practical guidance on choosing data generation strategies and highlight the distinct advantages and disadvantages of each approach.

动作识别虚拟数据IMU生成小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。