arXiv:2410.18099cs.CVcs.AI2024-10被引 1

用预训练+粗粒度轨迹编码,实现跨XR设备的高精度手势输入解码。

Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training

  • 通过粗粒度离散化轨迹并预训练,构建通用神经解码器。
  • 在4个数据集上平均Top-4准确率达90.4%,比SHARK²提升37.2%。
  • 模型仅4MB,实时运行仅需97毫秒,适合轻量化部署。

在扩展现实(XR)中,词组手势输入键盘(WGK)正成为主流交互方式。然而,不同交互模式、键盘尺寸和视觉反馈导致轨迹数据模式差异大,使解码复杂。现有模板匹配方法(如SHARK²)易受噪声影响;传统神经网络解码器虽更准,但需大量数据与深度学习知识。为此,我们提出一种结合易用性与高精度的新方案:基于大规模粗粒度离散轨迹预训练的通用神经解码器。该模型在增强现实(AR)与虚拟现实(VR)中的空中与表面式WGK系统间具有良好泛化能力,在四个多样化数据集上平均Top-4准确率达90.4%,显著优于SHARK²(提升37.2%),且超越传统神经解码器7.4%。模型经量化后仅4MB,可在Quest 3上以97毫秒延迟实现实时运行。

原文摘要 · Abstract (English)

Text entry with word-gesture keyboards (WGK) is emerging as a popular method and becoming a key interaction for Extended Reality (XR). However, the diversity of interaction modes, keyboard sizes, and visual feedback in these environments introduces divergent word-gesture trajectory data patterns, thus leading to complexity in decoding trajectories into text. Template-matching decoding methods, such as SHARK^2, are commonly used for these WGK systems because they are easy to implement and configure. However, these methods are susceptible to decoding inaccuracies for noisy trajectories. While conventional neural-network-based decoders (neural decoders) trained on word-gesture trajectory data have been proposed to improve accuracy, they have their own limitations: they require extensive data for training and deep-learning expertise for implementation. To address these challenges, we propose a novel solution that combines ease of implementation with high decoding accuracy: a generalizable neural decoder enabled by pre-training on large-scale coarsely discretized word-gesture trajectories. This approach produces a ready-to-use WGK decoder that is generalizable across mid-air and on-surface WGK systems in augmented reality (AR) and virtual reality (VR), which is evident by a robust average Top-4 accuracy of 90.4% on four diverse datasets. It significantly outperforms SHARK^2 with a 37.2% enhancement and surpasses the conventional neural decoder by 7.4%. Moreover, the Pre-trained Neural Decoder's size is only 4 MB after quantization, without sacrificing accuracy, and it can operate in real-time, executing in just 97 milliseconds on Quest 3.

手势输入XR交互神经解码轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。