arXiv:2511.01571cs.CVcs.RO2025-11被引 19

首个支持像素级理解的视觉语言动作模型,提升机器人操作精度与灵活性。

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

  • 采用多尺度像素感知编码器与视觉提示感知编码器融合架构。
  • 在3个基准上提升10.1%-28.7%操作成功率,仅需OpenVLA1.5%预训练成本。
  • 适合需要高精度场景理解的机器人控制研究者使用。

视觉-语言-动作模型(VLAs)正成为学习通用视觉运动控制策略的强大工具。然而,现有VLAs主要基于大规模图像-文本-动作数据训练,在像素级场景理解能力与依赖文本提示方面仍存在局限。为此,我们提出PixelVLA,首个支持像素级推理与多模态提示(文本+视觉)的VLA模型。其基于新型视觉运动指令微调框架,整合多尺度像素感知编码器与视觉提示感知编码器。为有效训练,我们进一步设计两阶段自动化标注流程,生成包含像素级标注的大规模数据集Pixel-160K,源自已有机器人数据。在三个标准VLA基准及两种VLA模型变体上的实验表明,PixelVLA相较OpenVLA将操作成功率提升10.1%-28.7%,且仅需其1.5%的预训练成本。结果证明,PixelVLA可无缝集成至现有VLAs中,实现复杂环境中更精准、高效、灵活的机器人控制。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image-text-action data and remain limited in two key ways: (i) they struggle with pixel-level scene understanding, and (ii) they rely heavily on textual prompts, which reduces their flexibility in real-world settings. To address these challenges, we introduce PixelVLA, the first VLA model designed to support both pixel-level reasoning and multimodal prompting with text and visual inputs. Our approach is built on a new visuomotor instruction tuning framework that integrates a multiscale pixel-aware encoder with a visual promptaware encoder. To train PixelVLA effectively, we further propose a two-stage automated annotation pipeline that generates Pixel-160K, a large-scale dataset with pixel-level annotations derived from existing robot data. Experiments on three standard VLA benchmarks and two VLA model variants show that PixelVLA improves manipulation success rates by 10.1%-28.7% over OpenVLA, while requiring only 1.5% of its pretraining cost. These results demonstrate that PixelVLA can be integrated into existing VLAs to enable more accurate, efficient, and versatile robot control in complex environments.

视觉语言动作像素级理解机器人控制多模态提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。