arXiv:2606.22480cs.RO2026-06被引 1

通过视觉对齐与迭代优化,提升机器人长程操作的量化技能精度。

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation

论文配图:ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation
图 1 · 摘自论文原文
  • 引入视觉-动作对齐损失,增强技能的视觉语义关联
  • 使用轻量级残差头实现两阶段精调,提升执行精度
  • 适用于需精准视觉引导的复杂机器人操作任务

长时序机器人操作的视觉-运动策略学习仍是核心挑战。基于离散量化技能的模仿学习方法虽有进展,但多仅将动作轨迹编码为隐变量技能,导致视觉语义关联弱,难以利用视觉观察进行技能选择;且离散分词不可避免引入连续动作生成的精度误差。为此,我们提出对齐精炼策略(ARP),一种耦合语义对齐与执行精炼的离散技能框架。具体包含:(i) 视觉-动作对齐目标,通过对比学习在共享潜空间中对齐视觉嵌入与预量化动作表示,同时保持状态无关的技能解码器;(ii) 轻量级迭代残差头(IRH),通过两步精调恢复细粒度控制以实现精确执行。大量实验表明,ARP在LIBERO和Meta-World基准上达到当前最优性能;真实机器人实验在Kuavo 4 Pro人形平台上验证其有效性,在两个挑战性操作任务中持续优于多个基线方法。

原文摘要 · Abstract (English)

Learning visuomotor policies for long-horizon manipulation remains a fundamental challenge. Recent skill-based imitation learning methods based on discrete quantization have shown promising results by representing complex behaviors as temporally extended skills. However, most existing approaches primarily encode action trajectories into latent skills, yielding weak visual-semantic grounding and limiting the ability to leverage visual observations for skill selection. Moreover, discrete tokenization inevitably incurs precision errors during continuous action generation. To alleviate these issues, we propose Aligned Refinement Policy (ARP), a discrete-skill framework that couples semantic grounding with execution-level refinement. Specifically, ARP introduces (i) a visual--action alignment objective that contrastively aligns visual embeddings with pre-quantized action representations in a shared latent space while preserving a state-independent skill decoder, and (ii) a lightweight Iterative Residual Head (IRH) that performs a two-step refinement to recover fine-grained control for precise execution. Extensive experiments show that ARP achieves state-of-the-art performance on the LIBERO and Meta-World benchmarks. Moreover, real-robot experiments on the Kuavo 4 Pro humanoid platform further validate its effectiveness, yielding consistent performance gains over several baselines on two challenging manipulation tasks.

机器人操作量化技能视觉对齐精炼控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。