arXiv:2602.20566cs.ROcs.CV2026-02被引 1

针对多视角机器人视觉模型,提出动态分层剪枝方法提升推理速度与操作成功率。

BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model

  • 分层设计:内视图预测保留关键区域,跨视图预测识别重要相机视角。
  • 在π0和RDT模型上成功率达10%提升,推理速度加快1.5倍至1.8倍。
  • 适合需要实时响应的机器人操作场景,尤其关注任务动态性与视图关联性。

视觉-语言-动作(VLA)模型通过融合大视觉语言模型(VLM)实现指令与视觉输入的联合理解,取得显著进展。然而,多视角输入带来的视觉令牌数量激增,严重制约了实时机器人操作的性能。现有VLM加速技术如令牌剪枝,在直接应用于VLA模型时表现下降,因其忽视不同视角间的关系,且未考虑机器人操作的任务动态特性。为此,本文提出BFA++,一种专为VLA模型设计的动态令牌剪枝框架。BFA++采用两级重要性预测器引导的分层剪枝策略:内视图预测器聚焦每张图像中的任务相关区域以抑制空间噪声,跨视图预测器识别不同操作阶段的关键摄像头以减少跨视角冗余。该设计在保留关键视觉线索的同时实现高效令牌选择,显著提升计算效率与操作成功率。在RoboTwin基准及真实机器人任务上的评估表明,BFA++在π0和RDT模型上分别实现约10%的成功率提升,推理速度分别提升1.8X和1.5X。结果表明,上下文敏感且任务感知的令牌剪枝比全视觉处理更有效,可推动实际机器人系统实现更快推理与更高精度。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. However, the substantial increase in visual tokens, particularly from multi-view inputs, poses serious challenges to real-time robotic manipulation. Existing acceleration techniques for VLMs, such as token pruning, often result in degraded performance when directly applied to VLA models, as they overlook the relationships between different views and fail to account for the dynamic and task-specific characteristics of robotic operation. To address this, we propose BFA++, a dynamic token pruning framework designed specifically for VLA models. BFA++ introduces a hierarchical pruning strategy guided by two-level importance predictors: an intra-view predictor highlights task-relevant regions within each image to suppress spatial noise, while an inter-view predictor identifies critical camera views throughout different manipulation phases to reduce cross-view redundancy. This design enables efficient token selection while preserving essential visual cues, resulting in improved computational efficiency and higher manipulation success rates. Evaluations on the RoboTwin benchmark and real-world robotic tasks demonstrate that BFA++ consistently outperforms existing methods. BFA++ improves the success rate by about 10% on both the π0 and RDT models, achieving speedup of 1.8X and 1.5X, respectively. Our results highlight that context-sensitive and task-aware token pruning serves as a more effective strategy than full visual processing, enabling faster inference and improved manipulation accuracy in real-world robotic systems.

视觉语言机器人令牌剪枝多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。