提出高效视觉语言动作模型,专攻实时机器人操作挑战
Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

- 通过未来预测与多帧融合增强时间推理能力
- 在动态任务中性能显著提升,延迟降低40%以上
- 适合需要快速反应的机器人应用场景
视觉-语言-动作(VLA)模型在机器人操作中表现优异,但现有基准主要评估静态任务泛化能力,忽视动态交互场景。为此,我们提出ReflexBench基准,包含六个动态任务,支持异步/同步推理下的可配置延迟,实现仿真步长与机器人控制解耦。基于此,我们设计了ReflexVLA模型,无需大规模机器人数据预训练即可实现反应关键型操作。该模型通过视觉骨干中的潜在未来预测和多帧时序融合增强时间推理,并利用批量视觉编码与CUDA图重放降低部署延迟。实验表明,ReflexVLA在动态任务中持续提升性能,同时在标准静态任务上保持竞争力;真实世界测试进一步验证其在实际部署条件下的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address this gap, we present ReflexBench, a benchmark for reaction-critical manipulation. ReflexBench contains six dynamic tasks and introduces an evaluation framework that decouples simulator stepping from robot control while supporting configurable latency under synchronous and asynchronous inference. Building upon ReflexBench, we propose ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining. ReflexVLA enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched visual encoding and CUDA Graph replay. Experiments show that ReflexVLA consistently improves dynamic manipulation performance while maintaining competitive accuracy on standard static manipulation benchmarks, and real-world experiments further demonstrate its effectiveness under practical deployment conditions. Project website: https://reflexvla.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。