arXiv:2607.08448cs.RO2026-07被引 8

用记忆代理让冻结的视觉语言模型更可靠地完成复杂操作。

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

论文配图:Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
图 1 · 摘自论文原文
  • 用记忆增强的智能体框架调用冻结的视觉语言模型,实现接触丰富的精准控制。
  • 在多个扰动场景下,相比最强基线提升38.6和25.4个百分点,成功率达58.4%。
  • 适合需要鲁棒语义理解与精细操作结合的机器人任务,如家庭环境操作。

语言驱动的操作需要精确的接触控制与对语言、场景和长时程的鲁棒推理能力。端到端的视觉-语言-动作(VLA)模型具备强大的局部视觉运动技能,但其训练依赖于分布内任务轨迹,在部署时易因语义重定向、目标重绑定、空间布局变化及不稳定接触而失败。大语言模型(LLM)编码智能体提供互补的语义与组合推理能力,但纯分析性原语难以应对不规则抓取、受限放置及刚性物体交互。本文提出Harness VLA:一种基于记忆增强的智能体框架,将冻结的VLA作为可重试的接触丰富原语,并与一组固定的小型分析性原语(用于定位、准备、运输、导航、释放)组合使用。无需扩展技能库,该框架通过任务特定执行轨迹、全局成功规则与失败模型学习这些固定原语的有效操作范围。通过将语义重定位、非接触执行与VLA重新调度交由规划器处理,同时保留冻结的VLA负责局部接触密集阶段,Harness VLA在不微调的前提下扩展了预训练VLA的适用范围。在受扰动的桌面、家庭厨房及从清洁到随机化双臂操作场景中,Harness VLA在LIBERO-Pro和RoboCasa365上分别优于最强基线38.6和25.4个百分点,在RoboTwin C2R上达到58.4%的成功率。

原文摘要 · Abstract (English)

Language-conditioned manipulation requires both precise contact-rich control and robust reasoning over language, scenes, and long horizons. End-to-end Vision-Language-Action (VLA) models provide strong local visuomotor skills, but they are trained on in-distribution task trajectories and often fail under deployment perturbations such as semantic retargeting, goal re-binding, spatial-layout shifts, and unstable local contacts. LLM coding agents provide complementary semantic and compositional reasoning, but purely analytic primitives struggle with irregular grasping, constrained placement, and articulated-object interaction. We present Harness VLA, a memory-augmented agentic framework that exposes a frozen VLA as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without finetuning. Across perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation, Harness VLA improves over the strongest relevant baselines by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and reaches 58.4% on RoboTwin C2R.

机器人操作视觉语言模型智能体系统泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。