arXiv:2607.04265cs.ROcs.AI2026-07被引 1

HALO-WA让机器人在真实操作中精准完成微调,成功率从26.4%提升至87.1%

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

论文配图:HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
图 1 · 摘自论文原文
  • 用混合注意力机制融合视觉和动作先验,动态优化生成动作序列
  • 实测四类任务成功率从26.4%提升至87.1%,优于最强基线19.2个百分点
  • 仅需45-75分钟在线训练,适合部署在需要高精度的机器人场景

世界-动作(WA)模型能生成适用于通用机器人操作的长时序动作块,但在真实场景中的校准、感知和接触动力学误差下仍易失效,尤其在对齐或插入的最后几毫米处失败。本文提出HALO-WA,一种基于混合注意力的潜空间引导在线强化学习框架,通过轻量级演员-评论家适配器利用WA生成过程中的潜特征与动作先验,实现快速在线适应真实部署误差。该框架引入混合注意力结构,在保持动作块时间一致性的同时,从受视觉上下文和末端修正需求条件化的WA潜空间中读取任务相关信息,从而生成优化后的动作块。我们在四个真实世界的精密操作任务上验证了HALO-WA,其平均成功率从基线模型的26.4%提升至87.1%,优于最强基线19.2个百分点,且每任务仅需45–75分钟在线训练。为促进可复现性,我们还在RoboTwin中补充了仿真实验,并公开代码于https://github.com/YeanRoot/HALO-WA。

原文摘要 · Abstract (English)

World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features and action priors from the WA generation process through a lightweight actor-critic adapter to enable fast online adaptation to real deployment errors. HALO-WA introduces a hybrid-attention structure that preserves the temporal consistency of action chunks while reading task-relevant information from WA latents conditioned on visual context and end-stage correction requirements, thereby producing refined action chunks. We validate HALO-WA on four real-world precision manipulation tasks, where it improves the average success rate from 26.4\% for WA-base to 87.1\%, outperforming the strongest baseline by 19.2 percentage points while requiring only 45--75 minutes of online training per task. To facilitate reproducibility, we further conduct supplementary simulation experiments in RoboTwin and release the code at https://github.com/YeanRoot/HALO-WA.

机器人操作强化学习动作生成在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。