评估音频图像视频弱监督方法在IMU动作定位中的适用性
WS-IMUBench: Can Weakly Supervised Methods from Audio, Image, and Video Be Adapted for IMU-based Temporal Action Localization?
- 用序列级标签测试7种弱监督方法在IMU数据上的表现
- 长动作和高维传感器数据下弱监督可达到较好效果
- 适合关注可扩展动作定位的学者与开发者
基于惯性测量单元(IMU)的人体活动识别已广泛应用于普适计算,但主流的片段分类范式难以捕捉真实行为的丰富时间结构。这推动了基于IMU的时间动作定位(IMU-TAL)的发展,即预测动作类别及其在连续流中的起止时间。然而,当前进展严重受限于密集的帧级边界标注,成本高且难扩展。为此,我们提出WS-IMUBench,一个在仅序列级标签下对弱监督IMU-TAL(WS-IMU-TAL)的系统性基准研究。不提出新算法,而是评估音频、图像、视频领域成熟的弱监督定位范式在仅序列标签下的迁移能力。我们在7个公开的IMU数据集上评测了7种代表性弱监督方法,完成超过3,540次模型训练与7,080次推理评估。基于三个研究问题:可迁移性、有效性与洞见,发现:(i) 迁移效果具有模态依赖性,时间域方法普遍比图像派生的提案方法更稳定;(ii) 在有利数据集上(如长动作、高维传感),弱监督可具竞争力;(iii) 主要失败模式来自短动作、时间模糊性和提案质量。最后,我们提出推进方向:如针对IMU的提案生成、边界感知目标与更强的时间推理。除具体结果外,WS-IMUBench建立了可复现的基准模板、数据集、协议与分析框架,助力社区加速可扩展的弱监督IMU-TAL发展。
原文摘要 · Abstract (English)
IMU-based Human Activity Recognition (HAR) has enabled a wide range of ubiquitous computing applications, yet its dominant clip classification paradigm cannot capture the rich temporal structure of real-world behaviors. This motivates a shift toward IMU Temporal Action Localization (IMU-TAL), which predicts both action categories and their start/end times in continuous streams. However, current progress is strongly bottlenecked by the need for dense, frame-level boundary annotations, which are costly and difficult to scale. To address this bottleneck, we introduce WS-IMUBench, a systematic benchmark study of weakly supervised IMU-TAL (WS-IMU-TAL) under only sequence-level labels. Rather than proposing a new localization algorithm, we evaluate how well established weakly supervised localization paradigms from audio, image, and video transfer to IMU-TAL under only sequence-level labels. We benchmark seven representative weakly supervised methods on seven public IMU datasets, resulting in over 3,540 model training runs and 7,080 inference evaluations. Guided by three research questions on transferability, effectiveness, and insights, our findings show that (i) transfer is modality-dependent, with temporal-domain methods generally more stable than image-derived proposal-based approaches; (ii) weak supervision can be competitive on favorable datasets (e.g., with longer actions and higher-dimensional sensing); and (iii) dominant failure modes arise from short actions, temporal ambiguity, and proposal quality. Finally, we outline concrete directions for advancing WS-IMU-TAL (e.g., IMU-specific proposal generation, boundary-aware objectives, and stronger temporal reasoning). Beyond individual results, WS-IMUBench establishes a reproducible benchmarking template, datasets, protocols, and analyses, to accelerate community-wide progress toward scalable WS-IMU-TAL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。