提出异步双频模型,让机器人更高效精准地执行语言指令。
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

- 分层设计:高层语义与底层控制解耦,提升计算效率。
- 在LIBERO-Plus上达85.5%成功率,动作生成提速近2.7倍。
- 适合需要快速响应的工业机器人任务,尤其关注跨平台泛化。
现实世界中视觉-语言-动作(VLA)模型的部署常受限于效率与性能的权衡、跨体态泛化能力以及执行平滑性。我们提出NebulaVLA,一种异步双频架构,将高层语义推理与底层动作控制解耦,优化计算资源并增强模块化。为弥合异构机器人间的语义鸿沟,引入GESTURE-7,一种统一的语言驱动动作表示。此外,我们的Guide Action算法通过基于掩码的平滑性约束,确保运动连续性。全面评估表明,NebulaVLA显著优于同步基线,在LIBERO-Plus上实现85.5%的平均成功率,并将动作生成速度提升约2.7倍。该异步设计使实际机器人具备高效且响应迅速的控制能力。
原文摘要 · Abstract (English)
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。