让视觉语言模型直接规划受力,无需训练就能完成复杂操作。
Unfettered Forceful Skill Acquisition with Physical Reasoning and Coordinate Frame Labeling
- 用坐标系标注增强视觉输入,让模型聚焦于受力而非路径。
- 4类任务220次实验中平均成功率51%,支持零样本泛化。
- 可自动修复失败操作,且能突破模型安全限制引发风险。
视觉语言模型(VLMs)具备丰富的物理世界知识,包括对物理属性、空间关系和运动的直觉理解。通过微调,它们可直接生成机器人轨迹。本文发现,通过引导模型输出作用力(wrench)而非轨迹,可显式激发其对力的推理能力,并在无预训练条件下实现零样本泛化。为此,在机器人摄像头图像上叠加一致的坐标系视觉标记作为查询增强。首先,该方法在四个任务(开/关盖、推杯/推椅)中验证,涵盖平移与旋转运动、力与位置量级差异、不同相机视角与标注方式,以及两台机器人平台,共220次实验,平均成功率达51%。其次,该框架使VLM能够持续根据交互反馈进行纠错,无论是否有监督。最后,我们观察到结合视觉标注与具身推理的提示策略可能绕过VLM的安全机制,分析了各提示组件对有害行为的贡献,讨论了具身推理发展的潜在风险。代码、视频及数据已公开:https://scalingforce.github.io/
原文摘要 · Abstract (English)
Vision language models (VLMs) exhibit vast knowledge of the physical world, including intuition of physical and spatial properties, affordances, and motion. With fine-tuning, VLMs can also natively produce robot trajectories. We demonstrate that eliciting wrenches, not trajectories, allows VLMs to explicitly reason about forces and leads to zero-shot generalization in a series of manipulation tasks without pretraining. We achieve this by overlaying a consistent visual representation of relevant coordinate frames on robot-attached camera images to augment our query. First, we show how this addition enables a versatile motion control framework evaluated across four tasks (opening and closing a lid, pushing a cup or chair) spanning prismatic and rotational motion, an order of force and position magnitude, different camera perspectives, annotation schemes, and two robot platforms over 220 experiments, resulting in 51% success across the four tasks. Then, we demonstrate that the proposed framework enables VLMs to continually reason about interaction feedback to recover from task failure or incompletion, with and without human supervision. Finally, we observe that prompting schemes with visual annotation and embodied reasoning can bypass VLM safeguards. We characterize prompt component contribution to harmful behavior elicitation and discuss its implications for developing embodied reasoning. Our code, videos, and data are available at: https://scalingforce.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。