arXiv:2604.17787cs.ROcs.AI2026-04被引 1

分步规划与修正动作,提升机器人操作精度

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models

论文配图:AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 将动作分为粗略轨迹和精细修正两阶段生成
  • 仿真成功率最高提升7.8%,真实场景提升18%
  • 适合需要高精度控制的机器人任务研究者

精密操控既需全局路径规划,又需执行中的局部修正,但现有视觉-语言-动作(VLA)模型通常在统一空间生成动作,导致大范围运动主导学习,抑制了对失败至关重要的微小修正信号。人类操作则依赖全局规划与执行中持续调整。为此,我们提出AnchorRefine,一种分层框架,将动作建模分解为轨迹锚点与残差修正两部分:锚点规划器生成粗略运动骨架,修正模块则针对执行偏差进行优化,提升几何与接触精度。我们还引入决策感知夹爪修正机制,更好捕捉夹爪控制的离散性和边界敏感性。在LIBERO、CALVIN及真实机器人任务上的实验表明,AnchorRefine可稳定提升基于回归与扩散的VLA骨干网络性能,仿真成功率最高提升7.8%,真实世界成功率提升18%。

原文摘要 · Abstract (English)

Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic formulation forces macro-level transport and micro-level refinement to be optimized under the same objective, causing large motions to dominate learning while suppressing small but failure-critical corrective signals. In contrast, human manipulation is structured by global movement planning together with continuous local adjustment during execution. Motivated by this principle, we propose AnchorRefine, a hierarchical framework that factorizes VLA action modeling into trajectory anchor and residual refinement. The anchor planner predicts a coarse motion scaffold, while the refinement module corrects execution-level deviations to improve geometric and contact precision. We further introduce a decision-aware gripper refinement mechanism to better capture the discrete and boundary-sensitive nature of gripper control. Experiments on LIBERO, CALVIN, and real-robot tasks demonstrate that AnchorRefine consistently improves both regression-based and diffusion-based VLA backbones, yielding gains of up to 7.8% in simulation success rate and 18% in real-world success rate.

机器人操作动作规划分层模型视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。