arXiv:2410.13605cs.LG2024-10被引 14

Transformer在传感器人体动作识别中表现不佳,因算力高、数据敏感且难部署。

Transformer-Based Approaches for Sensor-Based Human Activity Recognition: Opportunities and Challenges

  • 对比非Transformer方法,基于变换器的模型需更多计算资源
  • 在有限数据下性能更差,量化后准确率大幅下降
  • 对对抗攻击不鲁棒,影响实际应用可信度

Transformer在自然语言处理和计算机视觉中表现优异,正逐步应用于基于传感器的人体动作识别(HAR)。已有研究表明,只有在数据充足或使用高算力优化算法时,Transformer才能超越传统方法。然而,在传感器HAR领域,数据稀缺且常需在资源受限设备上训练与推理,这两种条件均难以满足。我们通过超过500次实验,系统比较了基于可穿戴传感器的Transformer与非Transformer方法,发现前者计算开销更大,性能普遍更低,量化后性能显著下降,且对对抗攻击更为脆弱,削弱用户对系统的信任。这些结果揭示了在真实场景中部署Transformer面临的严峻挑战。

原文摘要 · Abstract (English)

Transformers have excelled in natural language processing and computer vision, paving their way to sensor-based Human Activity Recognition (HAR). Previous studies show that transformers outperform their counterparts exclusively when they harness abundant data or employ compute-intensive optimization algorithms. However, neither of these scenarios is viable in sensor-based HAR due to the scarcity of data in this field and the frequent need to perform training and inference on resource-constrained devices. Our extensive investigation into various implementations of transformer-based versus non-transformer-based HAR using wearable sensors, encompassing more than 500 experiments, corroborates these concerns. We observe that transformer-based solutions pose higher computational demands, consistently yield inferior performance, and experience significant performance degradation when quantized to accommodate resource-constrained devices. Additionally, transformers demonstrate lower robustness to adversarial attacks, posing a potential threat to user trust in HAR.

人体动作识别Transformer资源受限模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。