arXiv:2603.22991cs.CV2026-03被引 1

不训练即可实现视觉标记剪枝,提升机器人操作效率

Training-Free Interaction-Aligned Visual Token Pruning for Efficient Embodied Manipulation

  • 通过语义与运动空间一致性动态调整每帧剪枝预算
  • 在真实机器人上实现1.48倍加速,仿真中达1.54倍速度提升
  • 适合追求高效视觉处理的机器人控制研究者

在具身操作中,策略需在闭环控制下反复处理密集的视觉标记序列,高效视觉表征是核心挑战。现有方法基于语义相关性、视觉语言模型注意力、跨帧冗余或动作空间中的运动进行标记排序或剪枝,但在指令相关的外观与观测图像运动尚未对齐时,可能丢弃任务相关区域。本文提出无需训练的交互对齐剪枝(IAprune),将每帧预算设定与预算内标记选择视为两个关联决策。语义-运动空间一致性决定采用保守或激进覆盖策略,所得区域大小映射为校准的动态预算。连续语义与运动响应对现有标记排序,几何残差修正则将固定选择槽位引导至欠代表的边界区域,不增加序列长度。在四个具身操作策略、三个仿真基准及一个真实机器人平台上,IAprune实现了优良的精度-效率权衡:在LIBERO上匹配未剪枝策略的同时达到1.54倍加速,在真实机器人上实现1.48倍加速。阶段分析表明,紧预算下早期阶段保留增益最大;固定预算分析确认,几何修正以边界和接触区域证据替代低优先级标记,而非简单保留更多标记。

原文摘要 · Abstract (English)

Efficient visual representation is a central image-processing challenge in embodied manipulation, where policies repeatedly process dense visual-token sequences during closed-loop control. Existing methods rank or prune tokens using semantic relevance, VLM attention, cross-frame redundancy, or motion in the action space. These signals may discard task-relevant regions when instruction-related appearance and observed image motion are not yet spatially aligned. We introduce Interaction-Aligned Pruning (IAprune), a training-free method that treats per-frame budget setting and within-budget token selection as two linked decisions. Semantic--motion spatial agreement guides the choice between Conservative and Aggressive coverage, and the resulting region size is mapped to a calibrated dynamic budget. Continuous semantic and motion responses rank the existing tokens, while geometric residual correction redirects fixed selection slots toward under-represented boundaries without increasing the sequence length. Across four embodied manipulation policies, three simulation benchmarks, and a real-robot platform, IAprune provides a favorable accuracy--efficiency trade-off, matching the unpruned policy on LIBERO with a \(1.54\times\) speedup and reaching \(1.48\times\) acceleration on a real robot. Phase-wise analysis shows that retention gains are largest under tight budgets early in an episode, while fixed-budget analysis confirms that geometry replaces low-priority tokens with boundary and contact-region evidence rather than retaining more tokens. Our project website is: \href{https://chengjt1999.github.io/VLA-IAP.github.io/}{IAprune.com}.

视觉剪枝机器人操作高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。