融合多模态意图识别与虚拟阻抗引导,提升虚拟现实遥操作抓取效率。
MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-based Multi-object Teleoperation
- 用多模态神经网络分析凝视、机械臂动作等信号,识别操作者意图。
- 虚拟阻抗模型使路径缩短30%以上,抓取成功率显著提升。
- 凝视数据最关键,适合远程操控、手术机器人等高精度场景。
虚拟现实(VR)中多物体遥操作的人机交互面临感知模糊和单模态意图识别局限的问题。本文提出一种共享控制框架,结合虚拟阻抗(VA)模型与基于多模态卷积神经网络的人类意图感知网络(MMIPN),以提升操作性能与用户体验。VA模型利用人工势场调节阻抗力,优化运动轨迹,引导操作者更高效地接近目标物。MMIPN融合凝视运动、机器人动作及环境上下文信息,估计人类抓取意图,缓解VR中的深度感知难题。用户实验对比四种条件,结果表明:MMIPN显著提高抓取成功率,VA模型使路径长度减少30%以上。凝视数据成为最关键的输入模态。研究验证了多模态线索与隐式引导结合在多物体抓取任务中的有效性,为未来多样化应用提供稳健解决方案。
原文摘要 · Abstract (English)
Effective human-robot interaction (HRI) in multi-object teleoperation tasks faces significant challenges due to perceptual ambiguities in virtual reality (VR) environments and the limitations of single-modality intention recognition. This paper proposes a shared control framework that combines a virtual admittance (VA) model with a Multimodal-CNN-based Human Intention Perception Network (MMIPN) to enhance teleoperation performance and user experience. The VA model employs artificial potential fields to guide operators toward target objects by adjusting admittance force and optimizing motion trajectories. MMIPN processes multimodal inputs, including gaze movement, robot motions, and environmental context, to estimate human grasping intentions, helping to overcome depth perception challenges in VR. Our user study evaluated four conditions across two factors, and the results showed that MMIPN significantly improved grasp success rates, while the VA model enhanced movement efficiency by reducing path lengths. Gaze data emerged as the most crucial input modality. These findings demonstrate the effectiveness of combining multimodal cues with implicit guidance in VR-based teleoperation, providing a robust solution for multi-object grasping tasks and enabling more natural interactions across various applications in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。