用视觉语言模型实现无障碍抓取,支持语音和手势指令。
OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
- 结合视觉、语音与开放词汇提示,实现多模态意图识别。
- 零样本检测未见过物体,抓取成功率高达87.00%。
- 适合残障人士在复杂环境中的自主抓取辅助使用。
抓取辅助对恢复运动障碍者在非结构化环境中的自主能力至关重要,因物体类别与用户意图多样且不可预测。我们提出OVGrasp,一种基于软外骨骼的分层控制框架,融合RGB-D视觉、开放词汇提示与语音指令,实现鲁棒的多模态交互。为提升开放环境下的泛化能力,OVGrasp采用视觉-语言基础模型与开放词汇机制,实现无需重训练的零样本对象检测。多模态决策器进一步融合空间与语言线索,推断多物体场景中的抓取或释放意图。我们在自研的头戴式视角可穿戴外骨骼上部署完整系统,并在15种物体上评估三种抓取类型。十名参与者实验表明,OVGrasp取得87.00%的抓取能力评分(GAS),优于现有最优基线,且运动学姿态更贴近自然手部动作。
原文摘要 · Abstract (English)
Grasping assistance is essential for restoring autonomy in individuals with motor impairments, particularly in unstructured environments where object categories and user intentions are diverse and unpredictable. We present OVGrasp, a hierarchical control framework for soft exoskeleton-based grasp assistance that integrates RGB-D vision, open-vocabulary prompts, and voice commands to enable robust multimodal interaction. To enhance generalization in open environments, OVGrasp incorporates a vision-language foundation model with an open-vocabulary mechanism, allowing zero-shot detection of previously unseen objects without retraining. A multimodal decision-maker further fuses spatial and linguistic cues to infer user intent, such as grasp or release, in multi-object scenarios. We deploy the complete framework on a custom egocentric-view wearable exoskeleton and conduct systematic evaluations on 15 objects across three grasp types. Experimental results with ten participants demonstrate that OVGrasp achieves a grasping ability score (GAS) of 87.00%, outperforming state-of-the-art baselines and achieving improved kinematic alignment with natural hand motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。