提出自监督方法联合建模手部与头部动作,提升XR交互体验
HaHeAE: Learning Generalisable Joint Representations of Human Hand and Head Movements in Extended Reality
- 基于图卷积与扩散模型的双编码器学习联合表征
- 重建质量提升74.0%,跨用户/场景泛化能力强
- 可支持动作聚类分析与生成,适配多种下游任务
手部和头部动作是扩展现实(XR)中最普遍的输入模态,对诸多应用具有重要意义。然而,以往研究多局限于单一模态或特定场景。本文提出HaHeAE——一种新型自监督方法,用于学习XR中手部与头部动作的通用联合表征。其核心为一个自编码器(AE),采用基于图卷积网络的语义编码器与基于扩散模型的随机编码器,分别学习动作的语义与随机表征,并通过扩散解码器重构原始信号。在三个公开XR数据集上的大量实验表明:1)重建质量相比常用自监督方法最高提升74.0%,且在不同用户、活动及XR环境中均具良好泛化性;2)支持可解释的手-头动作聚类识别与可变动作生成等新应用;3)可作为下游任务的有效特征提取器。结果证明该方法有效性,并凸显自监督联合建模手-头行为的潜力。
原文摘要 · Abstract (English)
Human hand and head movements are the most pervasive input modalities in extended reality (XR) and are significant for a wide range of applications. However, prior works on hand and head modelling in XR only explored a single modality or focused on specific applications. We present HaHeAE - a novel self-supervised method for learning generalisable joint representations of hand and head movements in XR. At the core of our method is an autoencoder (AE) that uses a graph convolutional network-based semantic encoder and a diffusion-based stochastic encoder to learn the joint semantic and stochastic representations of hand-head movements. It also features a diffusion-based decoder to reconstruct the original signals. Through extensive evaluations on three public XR datasets, we show that our method 1) significantly outperforms commonly used self-supervised methods by up to 74.0% in terms of reconstruction quality and is generalisable across users, activities, and XR environments, 2) enables new applications, including interpretable hand-head cluster identification and variable hand-head movement generation, and 3) can serve as an effective feature extractor for downstream tasks. Together, these results demonstrate the effectiveness of our method and underline the potential of self-supervised methods for jointly modelling hand-head behaviours in extended reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。