用大模型识别声学手势,让虚拟现实交互更自然高效
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
- 通过差分声学信道响应数据,结合大模型实现少样本手势识别
- 在10人、15种手势的实测数据上达到与传统方法相当的准确率
- 无需重新训练,适合快速部署于少样本虚拟现实场景
自然高效的交互仍是虚拟现实与增强现实系统的核心挑战。基于视觉的手势识别存在计算成本高、受光照影响大及隐私泄露等问题。声学传感提供了一种低成本、用户透明的替代方案:通过发射人耳不可闻的高频信号并捕捉其反射,通道冲激响应(CIR)编码了手势对声场的扰动。然而,现有基于CIR的手势识别方法通常依赖大规模标注数据进行模型训练,难以适用于少样本的VR场景。本文提出首个利用大语言模型(LLMs)进行VR/AR中基于CIR手势识别的框架。尽管大模型具备强大能力,但受限于手势特征不显著,实现少样本和零样本学习仍具挑战。为此,我们采集差分CIR而非原始数据,并构建了一个真实世界数据集,包含10名参与者完成15种手势(数字、字母、形状三类),每类重复10次。在该数据集上,采用大模型分类器进行大量实验,结果表明,本框架在无需领域特定再训练的情况下,实现了与经典机器学习基线相当的准确率。
原文摘要 · Abstract (English)
Natural and efficient interaction remains a critical challenge for virtual reality and augmented reality (VR/AR) systems. Vision-based gesture recognition suffers from high computational cost, sensitivity to lighting conditions, and privacy leakage concerns. Acoustic sensing provides an attractive alternative: by emitting inaudible high-frequency signals and capturing their reflections, channel impulse response (CIR) encodes how gestures perturb the acoustic field in a low-cost and user-transparent manner. However, existing CIR-based gesture recognition methods often rely on extensive training of models on large labeled datasets, making them unsuitable for few-shot VR scenarios. In this work, we propose the first framework that leverages large language models (LLMs) for CIR-based gesture recognition in VR/AR systems. Despite LLMs' strengths, it is non-trivial to achieve few-shot and zero-shot learning of CIR gestures due to their inconspicuous features. To tackle this challenge, we collect differential CIR rather than original CIR data. Moreover, we construct a real-world dataset collected from 10 participants performing 15 gestures across three categories (digits, letters, and shapes), with 10 repetitions each. We then conduct extensive experiments on this dataset using an LLM-adopted classifier. Results show that our LLM-based framework achieves accuracy comparable to classical machine learning baselines, while requiring no domain-specific retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。