arXiv:2512.00903cs.CVcs.RO2025-12被引 22

轻量级模型实现4D时空理解,推理速度提升18倍

SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead

  • 用4D视觉几何变换器+时序缓存提取时空特征
  • 在边缘设备上达大模型性能,速度更快18倍,内存减少12倍
  • 适合部署在资源受限的机器人场景

基于预训练视觉-语言模型(VLM)的视觉-语言-动作(VLA)模型潜力巨大,但因参数量过大而难以实用。虽已有轻量VLM方案,却牺牲了时空推理能力。现有方法虽引入3D输入改善表现,仍依赖大VLM融合多模态信息,且缺乏时间建模。为此,我们提出SwiftVLA:通过预训练的4D视觉几何变换器与时序缓存从2D图像中提取4D特征;引入可学习的融合令牌(Fusion Tokens),以未来预测目标训练生成统一的动作表示;并设计掩码重构策略,对VLM掩码4D输入并重建,使其学习有效4D表征,从而在推理时移除4D分支且性能损失极小。实测表明,SwiftVLA在真实与仿真环境中均超越轻量基线,媲美参数量高达7倍的VLA,在边缘设备上实现相当性能,速度提升18倍,内存降低12倍。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models built on pretrained Vision-Language Models (VLMs) show strong potential but are limited in practicality due to their large parameter counts. To mitigate this issue, using a lightweight VLM has been explored, but it compromises spatiotemporal reasoning. Although some methods suggest that incorporating additional 3D inputs can help, they usually rely on large VLMs to fuse 3D and 2D inputs and still lack temporal understanding. Therefore, we propose SwiftVLA, an architecture that enhances a compact model with 4D understanding while preserving design efficiency. Specifically, our approach features a pretrained 4D visual geometry transformer with a temporal cache that extracts 4D features from 2D images. Then, to enhance the VLM's ability to exploit both 2D images and 4D features, we introduce Fusion Tokens, a set of learnable tokens trained with a future prediction objective to generate unified representations for action generation. Finally, we introduce a mask-and-reconstruct strategy that masks 4D inputs to the VLM and trains the VLA to reconstruct them, enabling the VLM to learn effective 4D representations and allowing the 4D branch to be dropped at inference with minimal performance loss. Experiments in real and simulated environments show that SwiftVLA outperforms lightweight baselines and rivals VLAs up to 7 times larger, achieving comparable performance on edge devices while being 18 times faster and reducing memory footprint by 12 times.

轻量模型4D感知机器人控制边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。