用空间可操作性场引导视觉语言模型,让机器人摆脱记忆陷阱。
Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
- 引入可操作性场作为插件,动态引导机器人决策
- 实测在分布外场景下平均提升23.5%性能
- 适合需要鲁棒操作的现实机器人系统
视觉-语言-动作(VLA)模型在机器人操作中表现优异,能直接将视觉观测和语言指令映射为动作。然而,在分布外场景下仍显脆弱:测试环境变化时,VLA常重复记忆轨迹而非适应新场景,这种现象称为“记忆陷阱”。其根源在于端到端设计缺乏显式三维空间推理,难以在陌生环境中可靠识别可操作区域。为此,我们提出可操作性场干预(AFI)框架,利用3D空间可操作性场(SAFs)提供几何化可操作提示,明确标识机器人应接近或避开的区域。系统通过本体感知检测记忆陷阱,重定位至高可操作性区域,并生成基于可操作性的路径点,锚定VLA动作;再由基于SAF的评分器选择累积可操作性最高的轨迹。大量实验表明,该方法在真实机器人平台上的不同VLA骨干网络($π_{0}$ 和 $π_{0.5}$)下,分布外场景平均提升23.5%,在LIBERO-Pro基准上提升20.2%,验证了其增强VLA鲁棒性的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they remain brittle under distribution shifts: when test scenarios change, VLAs often reproduce memorized trajectories instead of adapting to the updated scene, which is a failure mode we refer to as the "Memory Trap". This limitation stems from the end-to-end design, which lacks explicit 3D spatial reasoning and prevents reliable identification of actionable regions in unfamiliar environments. To compensate for this missing spatial understanding, 3D Spatial Affordance Fields (SAFs) can provide a geometric representation that highlights where interactions are physically feasible, offering explicit cues about regions the robot should approach or avoid. We therefore introduce Affordance Field Intervention (AFI), a lightweight hybrid framework that uses SAFs as an on-demand plug-in to guide VLA behavior. Our system detects memory traps through proprioception, repositions the robot to recent high-affordance regions, and proposes affordance-driven waypoints that anchor VLA-generated actions. A SAF-based scorer then selects trajectories with the highest cumulative affordance. Extensive experiments demonstrate that our method achieves an average improvement of 23.5% across different VLA backbones ($π_{0}$ and $π_{0.5}$) under out-of-distribution scenarios on real-world robotic platforms, and 20.2% on the LIBERO-Pro benchmark, validating its effectiveness in enhancing VLA robustness to distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。