用头戴式陀螺仪识别行为,超越基础动作分类。
Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

- 构建分层模型,利用时间上下文提升行为识别能力。
- 在五类行为和八类场景上均超越现有方法,准确率达87.3%。
- 揭示哪些行为可稳定识别,适合智能眼镜实时辅助场景。
AR智能眼镜需要持续的行为上下文以提供主动帮助,但其最实用的始终开启传感器——头戴式惯性测量单元(IMU)——仅能检测行走、站立等基本运动模式。本文突破运动基元限制,定义了五个兼顾AR应用需求与传感器可观测性的行为类别。为此,我们构建了一个包含16万样本的Ego4D数据集,采用四级质量保障框架覆盖8个活动场景,并提出HiT-HAR模型(703K参数),在五类行为和八类场景识别任务上均优于现有头戴式IMU模型。通过每类行为的可区分性分析,揭示了部分行为(如移动)可稳定观测,部分需时间上下文(如物品传递、任务操作),而场景依赖信号重叠仍构成挑战。结果表明,利用时间上下文和场景结构的架构设计,优于单纯扩大模型规模。代码与数据集已开源。
原文摘要 · Abstract (English)
AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertial Measurement Unit (IMU), detects only motion primitives such as walking or standing. We push beyond motion primitives to behavioral-level recognition, defining five categories that balance AR application need with sensor observability. To this end, we construct a 160K-sample Ego4D dataset with a four-tier quality assurance framework spanning 8 activity scenarios, and propose HiT-HAR, a 703K-parameter hierarchical model that outperforms prior head-mounted IMU models on five-class action and eight-class scenario recognition. We further map the observability frontier of head-mounted IMU through per-class separability analysis, identifying which behavioral categories are reliably observable (Locomotion), which benefit from temporal context (Object Transfer, Task Operation), and where scenario-dependent signal overlap poses remaining challenges. Our results indicate that architectural choices exploiting temporal context and scenario structure outperform simply scaling model size. The code and dataset are publicly available at https://github.com/Harvard-AI-and-Robotics-Lab/HiT-HAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。