arXiv:2509.20579cs.CVcs.LG2025-09中稿 · 2025 IEEE-RAS 24th…被引 1

用视觉模型注意力图提升机器人双手操作的3D感知能力

Large Pre-Trained Models for Bimanual Manipulation in 3D

  • 将DINOv2的注意力图转为3D体素语义线索
  • 在RLBench双臂任务中平均提升8.2%,相对增益21.9%
  • 适合关注视觉引导机器人控制的研究者

我们研究将预训练视觉变换器(Vision Transformer)的注意力图融入体素表示,以增强双臂机器人操作能力。具体而言,从自监督的ViT模型DINOv2中提取注意力图,并将其解释为RGB图像的像素级显著性评分。这些图被映射到3D体素网格中,生成体素级语义线索,并融入行为克隆策略。当集成到最先进的基于体素的策略中时,该方法在RLBench双臂基准测试的所有任务上实现平均绝对提升8.2%和相对增益21.9%。

原文摘要 · Abstract (English)

We investigate the integration of attention maps from a pre-trained Vision Transformer into voxel representations to enhance bimanual robotic manipulation. Specifically, we extract attention maps from DINOv2, a self-supervised ViT model, and interpret them as pixel-level saliency scores over RGB images. These maps are lifted into a 3D voxel grid, resulting in voxel-level semantic cues that are incorporated into a behavior cloning policy. When integrated into a state-of-the-art voxel-based policy, our attention-guided featurization yields an average absolute improvement of 8.2% and a relative gain of 21.9% across all tasks in the RLBench bimanual benchmark.

双臂操作视觉注意体素表示机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。