arXiv:2410.08779cs.CVcs.HC2024-10被引 1

用视觉模型将手势映射到二维空间,实现稳定流畅的空中手势交互。

HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive Systems

  • 基于变分自编码器将手部姿态映射到二维可视空间。
  • 在12人实验中,使用手势引导窗口可显著减少任务完成时间和目标距离。
  • 通过数据增强与滤波优化,提升手势交互的稳定性与平滑性,适合多媒体探索场景。

我们提出一种新型的手部姿态嵌入交互系统(HpEIS),作为一种虚拟传感器,利用在多种手部姿态上训练的变分自编码器(VAE),将用户灵活的手势映射到二维视觉空间。该系统仅需摄像头即可实现对多媒体资源的可视化、可引导的探索交互。通过与专家及非专业用户的初步实验,识别出系统稳定性与平滑性方面的通用问题。为此,我们设计了多项改进:手部姿态数据增强、损失函数中加入抗抖动正则项、运动转折点的稳定后处理,以及基于One Euro滤波器的平滑后处理。在目标选择实验中(n=12),通过任务完成时间与最终距离目标点的距离评估了有无手势引导窗口的效果。实验结果表明,HpEIS为用户提供了一种可学习、灵活、稳定且平滑的空中手势交互体验。

原文摘要 · Abstract (English)

We present a novel Hand-pose Embedding Interactive System (HpEIS) as a virtual sensor, which maps users' flexible hand poses to a two-dimensional visual space using a Variational Autoencoder (VAE) trained on a variety of hand poses. HpEIS enables visually interpretable and guidable support for user explorations in multimedia collections, using only a camera as an external hand pose acquisition device. We identify general usability issues associated with system stability and smoothing requirements through pilot experiments with expert and inexperienced users. We then design stability and smoothing improvements, including hand-pose data augmentation, an anti-jitter regularisation term added to loss function, stabilising post-processing for movement turning points and smoothing post-processing based on One Euro Filters. In target selection experiments (n=12), we evaluate HpEIS by measures of task completion time and the final distance to target points, with and without the gesture guidance window condition. Experimental responses indicate that HpEIS provides users with a learnable, flexible, stable and smooth mid-air hand movement interaction experience.

手势识别交互系统三维交互视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。