arXiv:2410.17489cs.CVcs.AI2024-10中稿 · the Proceedings of…被引 3

无需标注数据,通过自集成与分布对齐提升动作识别泛化能力

Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment

  • 用时序集成生成稳定伪标签,减少噪声干扰
  • 通过一致性正则化增强数据增广效果,提升小样本泛化性能
  • 适合用户差异大、标注稀缺的动作识别场景

基于深度学习的可穿戴人体动作识别(wHAR)虽提升了复杂动作捕捉与分类能力,但受限于专家标注缺失和用户间域差异,泛化能力不足。数据增强可改善泛化性,但无监督增广需谨慎以避免引入噪声。无监督域适应(UDA)通过对齐条件分布缓解域差异,但传统伪标签易引发误差传播。本文提出μDAR,一种联合优化架构,包含三部分:(i) 增强样本间一致性正则化以提升分类泛化性;(ii) 时序集成用于鲁棒伪标签生成;(iii) 条件分布对齐,通过最小化源-目标特征空间间的核类条件最大均值差异(kCMMD)学习域不变嵌入。时序集成聚合历史预测,平滑伪标签噪声,再用于条件分布对齐。一致性正则化确保同一样本多轮增广共享相同标签,从而在仅有限源数据下实现强泛化,并保证目标样本伪标签一致性。μDAR在四个基准wHAR数据集上相较六种先进UDA方法平均宏F1提升约4-12%。

原文摘要 · Abstract (English)

Recent advancements in deep learning-based wearable human action recognition (wHAR) have improved the capture and classification of complex motions, but adoption remains limited due to the lack of expert annotations and domain discrepancies from user variations. Limited annotations hinder the model's ability to generalize to out-of-distribution samples. While data augmentation can improve generalizability, unsupervised augmentation techniques must be applied carefully to avoid introducing noise. Unsupervised domain adaptation (UDA) addresses domain discrepancies by aligning conditional distributions with labeled target samples, but vanilla pseudo-labeling can lead to error propagation. To address these challenges, we propose $μ$DAR, a novel joint optimization architecture comprised of three functions: (i) consistency regularizer between augmented samples to improve model classification generalizability, (ii) temporal ensemble for robust pseudo-label generation and (iii) conditional distribution alignment to improve domain generalizability. The temporal ensemble works by aggregating predictions from past epochs to smooth out noisy pseudo-label predictions, which are then used in the conditional distribution alignment module to minimize kernel-based class-wise conditional maximum mean discrepancy ($k$CMMD) between the source and target feature space to learn a domain invariant embedding. The consistency-regularized augmentations ensure that multiple augmentations of the same sample share the same labels; this results in (a) strong generalization with limited source domain samples and (b) consistent pseudo-label generation in target samples. The novel integration of these three modules in $μ$DAR results in a range of $\approx$ 4-12% average macro-F1 score improvement over six state-of-the-art UDA methods in four benchmark wHAR datasets

动作识别无监督域适应自集成穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。