arXiv:2508.10333cs.ROcs.CV2025-08AAAI被引 86

通过重建视觉注意力目标,让机器人更精准地识别并操作物体。

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

  • 用扩散变换器重建注视区域,隐式引导视觉注意力聚焦
  • 在100k轨迹、200万样本数据上训练,提升泛化能力
  • 实测证明在仿真和真实世界均实现高精度操作

近期视觉-语言-动作(VLA)模型虽推动了机器人多模态理解与动作执行的融合,但实证分析发现现有模型难以将视觉注意力集中于目标区域,常呈现分散状态。为此,本文提出ReconVLA,一种基于隐式对齐范式的重构型VLA模型。该模型以视觉输出为条件,通过扩散变换器重建图像中的注视区域,对应被操作物体的目标位置。这一机制促使模型学习细粒度表征并精准分配视觉注意力,有效利用任务相关的视觉信息,实现精确操控。此外,我们构建了一个大规模预训练数据集,涵盖超过100,000条轨迹和200万条数据样本,来自开源机器人数据集,进一步增强模型在视觉重建上的泛化能力。大量仿真与真实世界实验验证了该方法在精确操作和跨场景泛化方面的优越性。

原文摘要 · Abstract (English)

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispersed. To guide the visual attention grounding on the correct target, we propose ReconVLA, a reconstructive VLA model with an implicit grounding paradigm. Conditioned on the model's visual outputs, a diffusion transformer aims to reconstruct the gaze region of the image, which corresponds to the target manipulated objects. This process prompts the VLA model to learn fine-grained representations and accurately allocate visual attention, thus effectively leveraging task-specific visual information and conducting precise manipulation. Moreover, we curate a large-scale pretraining dataset comprising over 100k trajectories and 2 million data samples from open-source robotic datasets, further boosting the model's generalization in visual reconstruction. Extensive experiments in simulation and the real world demonstrate the superiority of our implicit grounding method, showcasing its capabilities of precise manipulation and generalization. Our project page is https://zionchow.github.io/ReconVLA/.

机器人感知视觉注意力多模态模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。