让机器人通过局部视觉线索快速学会精细操作,提升成功率32.5%。
GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

- 用区域级监督学习视角相关的视觉锚点,聚焦关键局部特征。
- 在四类生物医学任务中,成功率提升32.5个百分点。
- 适合需要高效在线适应的精密机器人操作场景。
预训练的视觉-语言-动作(VLA)策略为机器人操作提供了强大先验,但将其在线适应到精细的生物医学任务仍具挑战性。任务成功常依赖细微且视角相关的视觉线索,而任务级奖励对哪些区域重要缺乏指导,导致在有限真实机器人交互下难以学习任务相关的视觉定位。在线适应还受限于VLA推理的计算成本和基于回放的更新开销。我们提出GRAFT(Grounded Reinforcement Adaptation for Fast Task Learning),一种通过视觉定位实现高效在线VLA适应的框架。GRAFT利用区域级监督学习视图特定的视觉锚点,聚焦任务相关局部线索,无需部署时生成区域提议。同时结合单步动作生成与缓存的视觉-语言前缀复用,加速在线学习。在四个生物医学操作任务中,GRAFT在相同适应预算下成功率提升32.5个百分点,且显著降低在线策略更新的计算开销。
原文摘要 · Abstract (English)
Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-dependent visual cues, while task-level rewards provide little guidance about which regions matter, making it difficult to learn task-relevant visual grounding from limited real-robot interaction. Online adaptation is further constrained by the computational cost of VLA inference and replay-based updates. We introduce GRAFT (Grounded Reinforcement Adaptation for Fast Task Learning), a framework for efficient online VLA adaptation through grounded perception. GRAFT uses region-level supervision to learn view-specific visual anchors that focus perception on task-relevant local cues without requiring region proposals at deployment. It further combines single-step action generation with cached visual-language prefix reuse to accelerate online learning. Across four biomedical manipulation tasks, GRAFT improves success rates by 32.5 percentage points under matched adaptation budgets, while reducing the computational overhead of online policy updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。