arXiv:2509.01031cs.LGcs.AI2025-09

用强化学习提取跨用户通用动作特征,无需用户标注即可提升识别准确率。

Reinforcement Learning Driven Generalizable Feature Representation for Cross-User Activity Recognition

  • 用强化学习将特征提取变为序列决策过程,提升跨用户泛化能力。
  • 在DSADS和PAMAP2数据集上准确率超越现有方法,无需为新用户校准。
  • 无需目标用户标签,自动学习时间一致的通用动作模式,适合实际部署。

基于可穿戴传感器的人体动作识别(HAR)对医疗健康、健身追踪和智能环境至关重要,但用户间差异——包括运动模式、传感器位置和生理特征不同——严重阻碍了真实场景中的泛化性能。传统监督学习易过拟合于特定用户,导致对未见用户的识别效果差。现有领域泛化方法虽有潜力,却常忽略时序依赖性或依赖不切实际的领域标签。本文提出时序保持强化学习领域泛化框架(TPRL-DG),将特征提取重构为由强化学习驱动的序列决策过程。该框架采用基于Transformer的自回归生成器,生成捕捉用户无关动作动态的时间标记,并通过多目标奖励函数优化,兼顾类别区分与跨用户不变性。核心创新包括:(1) 强化学习驱动的领域泛化方法;(2) 自回归标记化以保持时序一致性;(3) 无标签奖励设计,无需目标用户标注。在DSADS和PAMAP2数据集上的评估表明,TPRL-DG在跨用户泛化方面优于当前最优方法,实现更高准确率且无需逐用户校准。通过学习鲁棒的、用户无关的时序模式,该方法推动了可扩展的HAR系统发展,助力个性化医疗、自适应健身追踪与情境感知环境的落地。

原文摘要 · Abstract (English)

Human Activity Recognition (HAR) using wearable sensors is crucial for healthcare, fitness tracking, and smart environments, yet cross-user variability -- stemming from diverse motion patterns, sensor placements, and physiological traits -- hampers generalization in real-world settings. Conventional supervised learning methods often overfit to user-specific patterns, leading to poor performance on unseen users. Existing domain generalization approaches, while promising, frequently overlook temporal dependencies or depend on impractical domain-specific labels. We propose Temporal-Preserving Reinforcement Learning Domain Generalization (TPRL-DG), a novel framework that redefines feature extraction as a sequential decision-making process driven by reinforcement learning. TPRL-DG leverages a Transformer-based autoregressive generator to produce temporal tokens that capture user-invariant activity dynamics, optimized via a multi-objective reward function balancing class discrimination and cross-user invariance. Key innovations include: (1) an RL-driven approach for domain generalization, (2) autoregressive tokenization to preserve temporal coherence, and (3) a label-free reward design eliminating the need for target user annotations. Evaluations on the DSADS and PAMAP2 datasets show that TPRL-DG surpasses state-of-the-art methods in cross-user generalization, achieving superior accuracy without per-user calibration. By learning robust, user-invariant temporal patterns, TPRL-DG enables scalable HAR systems, facilitating advancements in personalized healthcare, adaptive fitness tracking, and context-aware environments.

动作识别强化学习跨用户泛化可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。