arXiv:2503.13542cs.LGcs.AI2025-03被引 5

用数据混合优化提升跨数据集动作识别的泛化能力。

HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets

  • 设计新框架,用均方误差替代语言模型损失,适配惯性传感器连续多通道数据。
  • 融合马洪尼算法校正不同设备朝向差异,实现坐标系对齐。
  • 仅用30%-50%数据就比现有方法平均高6.51%准确率,适合资源受限场景。

跨数据集人体动作识别因模型泛化能力有限而难以实际部署。受大语言模型中DoReMi成功的启发,本文提出一种面向自监督动作识别的数据混合优化策略,旨在提升异构惯性传感器数据下的识别性能。由于惯性测量单元(IMU)数据具有连续、多通道及内在异质性特征,直接应用传统DoReMi面临挑战。为此,我们提出HAR-DoReMi框架,引入基于均方误差(MSE)的掩码重建任务,取代原DoReMi中依赖负对数似然(NLL)的离散序列预测任务,更契合IMU数据特性。同时,该框架整合了马洪尼(Mahony)融合算法,在自监督预训练中估计各数据集内传感器朝向,实现与统一坐标系对齐,缓解设备朝向差异带来的异质性。在多个跨数据集迁移任务上的实验表明,与当前最优方法相比,HAR-DoReMi仅需约30%至50%的数据量,平均准确率提升6.51%,验证了其在提升泛化能力与数据效率方面的有效性,显著推动了动作识别技术的实际应用。

原文摘要 · Abstract (English)

Cross-dataset Human Activity Recognition (HAR) suffers from limited model generalization, hindering its practical deployment. To address this critical challenge, inspired by the success of DoReMi in Large Language Models (LLMs), we introduce a data mixture optimization strategy for pre-training HAR models, aiming to improve the recognition performance across heterogeneous datasets. However, directly applying DoReMi to the HAR field encounters new challenges due to the continuous, multi-channel and intrinsic heterogeneous characteristics of IMU sensor data. To overcome these limitations, we propose a novel framework HAR-DoReMi, which introduces a masked reconstruction task based on Mean Squared Error (MSE) loss. By raplacing the discrete language sequence prediction task, which relies on the Negative Log-Likelihood (NLL) loss, in the original DoReMi framework, the proposed framework is inherently more appropriate for handling the continuous and multi-channel characteristics of IMU data. In addition, HAR-DoReMi integrates the Mahony fusion algorithm into the self-supervised HAR pre-training, aiming to mitigate the heterogeneity of varying sensor orientation. This is achieved by estimating the sensor orientation within each dataset and facilitating alignment with a unified coordinate system, thereby improving the cross-dataset generalization ability of the HAR model. Experimental evaluation on multiple cross-dataset HAR transfer tasks demonstrates that HAR-DoReMi improves the accuracy by an average of 6.51%, compared to the current state-of-the-art method with only approximately 30% to 50% of the data usage. These results confirm the effectiveness of HAR-DoReMi in improving the generalization and data efficiency of pre-training HAR models, underscoring its significant potential to facilitate the practical deployment of HAR technology.

动作识别自监督学习多模态融合数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。