arXiv:2607.27180cs.CVcs.RO2026-07被引 1

让视觉语言模型在真实身体中执行任务,测试其具身智能水平。

HumanCLAW: Can Vision-Language Models Act Through a Body?

论文配图:HumanCLAW: Can Vision-Language Models Act Through a Body?
图 1 · 摘自论文原文
  • 分离决策与执行,用物理身体真实动作评估模型选择能力。
  • 9个顶尖模型在1218个任务中最高仅16.8%成功率,普遍缺乏身体感知。
  • 适合研究具身认知、智能体行为评估的学者使用。

评估视觉语言模型(VLM)是否能通过物理身体行动极具挑战性,因动作结果同时受模型决策与运动控制影响。任务失败时难以判断是模型决策失误还是执行失败(如失衡摔倒)。本文提出HumanCLAW框架,将动作决策与底层执行解耦:每一步由现成的VLM发出原子技能指令,再转化为包含重力与碰撞等物理效应的亚秒级全身连续动作。身体可自由在真实世界活动,而执行扰动、平衡与电机误差被剔除。唯一可测量的是模型的动作智能——即时决定下一步应执行的动作。基于此,构建了HumanCLAW-Bench:覆盖41个室内场景的1,218个长时程、第一人称“寻找-导航-交互”任务。测试9个前沿VLM后发现,无一能完成基准任务,最佳模型仅达16.8%成功率。识别目标并非瓶颈,当前VLM缺乏具身自知:无法追踪自身位置、是否抵达目标或是否碰撞障碍物。

原文摘要 · Abstract (English)

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

具身智能视觉语言模型动作评估人体模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。