arXiv:2602.18813cs.ROcs.LG2026-02被引 1

Habilis-β在设备端实现高速长时视觉语言动作控制,性能远超现有模型。

Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model

  • 融合无语言预训练与循环任务微调,构建强鲁棒交互先验
  • 1小时连续运行下达572.6任务/小时,平均干预间隔39.2秒
  • 适合真实机器人长期自主作业场景,尤其擅长快速复杂操作

我们提出Habilis-β,一种面向实际部署的高速、长时在设备端视觉-语言-动作(VLA)模型。现有VLA评估多局限于受控重置下的单次成功概率,无法反映真实场景所需的高速与持续性。为此,我们引入生产力-可靠性平面(PRP),通过任务/小时(TPH)和平均干预间隔(MTBI)在连续运行协议下评估性能。Habilis-β通过大规模游戏数据进行无语言预训练以获得稳健交互先验,并在循环任务演示上后训练以捕捉状态漂移。系统采用ESPADA实现相位自适应运动调节以加速自由空间移动,利用修正流蒸馏实现在边缘设备上的高频控制,结合无分类器引导(CFG)作为部署时可调参数,动态平衡指令遵循与学习先验。在1小时连续运行中,仿真环境下达到572.6 TPH和39.2秒MTBI(对比π₀.₅为120.5 TPH和30.5秒),真实人形物流任务中实现124 TPH和137.4秒MTBI(对比π₀.₅为19 TPH和46.1秒)。此外,Habilis-β在RoboTwin 2.0标准榜单上多项任务取得最高表现,验证其在复杂操作中的有效性。

原文摘要 · Abstract (English)

We introduce Habilis-$β$, a fast-motion and long-lasting on-device vision-language-action (VLA) model designed for real-world deployment. Current VLA evaluation remains largely confined to single-trial success rates under curated resets, which fails to capture the fast-motion and long-lasting capabilities essential for practical operation. To address this, we introduce the Productivity-Reliability Plane (PRP), which evaluates performance through Tasks per Hour (TPH) and Mean Time Between Intervention (MTBI) under a continuous-run protocol that demands both high-speed execution and sustained robustness. Habilis-$β$ achieves high performance by integrating language-free pre-training on large-scale play data for robust interaction priors with post-training on cyclic task demonstrations that capture state drift across consecutive task iterations. The system further employs ESPADA for phase-adaptive motion shaping to accelerate free-space transit, utilizes rectified-flow distillation to enable high-frequency control on edge devices, and incorporates classifier-free guidance (CFG) as a deployment-time knob to dynamically balance instruction adherence and learned interaction priors. In 1-hour continuous-run evaluations, Habilis-$β$ achieves strong performance under the PRP metrics, compared to $π_{0.5}$ in both simulation and real-world environments. In simulation, Habilis-$β$ achieves 572.6 TPH and 39.2 s MTBI (vs. 120.5 TPH and 30.5 s for $π_{0.5}$), while in a real-world humanoid logistics workflow it achieves 124 TPH and 137.4 s MTBI (vs. 19 TPH and 46.1 s for $π_{0.5}$). Finally, Habilis-$β$ achieves the highest reported performance on the standard RoboTwin 2.0 leaderboard across representative tasks, validating its effectiveness in complex manipulation scenarios.

视觉语言动作边缘计算机器人控制长时运行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。