用视觉语言模型理解用户意图,让机器人更聪明地辅助远程操控。
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
- 利用预训练视觉语言模型进行实时意图推理,结合常识知识理解操作输入。
- 在多种新场景和任务中表现优于基线,提升任务成功率并降低用户负担。
- 适合需要灵活协作的复杂远程操控场景,如救援或家庭服务。
辅助远程操控系统通过人机共享控制,在多样化且非结构化环境中实现高效直观的人机协作。其核心挑战在于机器人需从用户操控输入中推断广泛的人类意图,并采取正确协助动作。现有方法受限于简单预设场景或特定任务数据分布,难以支持真实世界应用。本文提出Casper系统,利用预训练视觉语言模型(VLMs)嵌入的常识知识,实现实时意图推断与灵活技能执行。Casper包含开放世界感知模块以泛化理解新物体与场景,基于VLM的意图推理机制利用常识推理解析用户操控片段,并构建技能库扩展了先前系统的任务覆盖范围,支持多样、长时程的移动操作任务。大量实证评估(包括人类实验与系统消融研究)表明,相较于直接遥控和现有辅助基线,Casper显著提升任务性能,减轻用户认知负荷,获得更高满意度。
原文摘要 · Abstract (English)
Assistive teleoperation, where control is shared between a human and a robot, enables efficient and intuitive human-robot collaboration in diverse and unstructured environments. A central challenge in real-world assistive teleoperation is for the robot to infer a wide range of human intentions from user control inputs and to assist users with correct actions. Existing methods are either confined to simple, predefined scenarios or restricted to task-specific data distributions at training, limiting their support for real-world assistance. We introduce Casper, an assistive teleoperation system that leverages commonsense knowledge embedded in pre-trained visual language models (VLMs) for real-time intent inference and flexible skill execution. Casper incorporates an open-world perception module for a generalized understanding of novel objects and scenes, a VLM-powered intent inference mechanism that leverages commonsense reasoning to interpret snippets of teleoperated user input, and a skill library that expands the scope of prior assistive teleoperation systems to support diverse, long-horizon mobile manipulation tasks. Extensive empirical evaluation, including human studies and system ablations, demonstrates that Casper improves task performance, reduces human cognitive load, and achieves higher user satisfaction than direct teleoperation and assistive teleoperation baselines. More information is available at https://ut-austin-rpl.github.io/casper/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。