arXiv:2604.00767cs.LG2026-04

提出开放叙事框架,让可穿戴设备理解连续、自定义的日常活动。

ActivityNarrated: An Open-Ended Narrative Paradigm for Wearable Human Activity Understanding

论文配图:ActivityNarrated: An Open-Ended Narrative Paradigm for Wearable Human Activity Understanding
图 1 · 摘自论文原文
  • 将连续动作转为可复用的动作标记,结合语言模型生成开放词汇描述。
  • 在长时序任务中实现31.6%的宏平均F1提升,优于现有最先进方法。
  • 适合需要理解复杂、个性化动作场景的研究与应用,如健康监护。

可穿戴人体活动识别(HAR)虽持续进步,但多数研究仍基于固定窗口、封闭集分类范式,难以匹配日常行为的开放性、无剧本性、个性化、时长可变及组合特性。为此,本文提出ActivityNarrated——一种语言引导的可穿戴活动理解开放叙事范式。将其建模为密集传感器信号字幕生成,并构建全面基准协议,评估时间定位精度、字幕质量、传感器-语言对齐度、传统封闭集分类表现及鲁棒性。进一步提出三阶段架构ActNarrator:将连续惯性测量单元(IMU)信号离散化为可复用动作标记,利用外部冻结的小型语言模型生成开放词汇活动描述。实验表明,该方法实现高质量的密集信号字幕生成,具备更强适应性与鲁棒性,可将基于传感器的活动理解转化为文本级推理,支持下游分类任务,在宏平均F1上超越当前最优模型3.8%-31.6%;同时支持跨长时间跨度的复杂问答,拓展全新理解能力。

原文摘要 · Abstract (English)

Wearable human activity recognition (HAR) has made steady progress, yet much of this progress remains grounded in fixed-window, closed-set classification benchmarks. This formulation is poorly matched to everyday behavior, where activities are open-ended, unscripted, personalized, variable in duration, and often compositional. To address this mismatch, we introduce ActivityNarrated, an open-ended narrative paradigm for language-grounded wearable activity understanding. We formulate this setting as dense sensor signal captioning with a comprehensive benchmark protocol that measures temporal localization, caption quality, sensor-language alignment, conventional closed-set classification as a downstream diagnostic, and additional robustness measures. We further present ActNarrator, a 3-stage architecture that discretizes continuous IMU signals into reusable motion tokens and uses an external frozen small language model to generate open-vocabulary activity captions. Experiments show that our method provides high quality dense sensor captioning with superior adaptivity and robustness, enabling various downstream tasks by turning sensor-based human activity understanding into sensor-grounded text-level reasoning. This includes downstream classification where ActNarrator outperforms state-of-the-art HAR models by 3.8 - 31.6 \% in Macro-F1. This paradigm also enables novel activity understanding capabilities such as complex question-answering over long time horizons.

活动识别语言模型可穿戴设备开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。