arXiv:2606.09572cs.ROcs.AI2026-06

轻量级视觉动作模型实现高效机器人控制,支持云端语义推理与本地实时执行。

CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control

论文配图:CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control
图 1 · 摘自论文原文
  • 受小脑-丘脑启发,分路注意力融合视觉、本体感觉与任务条件
  • 仅68M参数即达大型模型水平,在LIBERO上表现优异且延迟更低
  • 适合资源受限的机器人平台,支持高频闭环控制与异步执行

视觉-语言-动作模型在机器人操作中展现出巨大潜力,但通常需持续处理原始语言以表达任务意图,这在高频低层执行阶段效率低下。为解决这一问题,本文提出一种受小脑-丘脑启发的视觉动作模型(CT-VAM),用于高效的任务条件化视觉运动控制。CT-VAM作为紧凑的局部执行策略,从双视角视觉观测、本体感觉和轻量任务条件中预测动作块,支持云端大模型进行高层语义推理、本地硬件实现快速闭环控制的云-边协同范式。为有效融合异构输入,模型引入TARS(丘脑动作路由流),通过分路条件注意力解码器独立路由动作、视觉与任务流,避免密集感官令牌淹没关键任务信息。仅含68M参数的CT-VAM在LIBERO基准上取得与更大模型相当的成功率,同时显著降低推理延迟。结合流动一致的图像修复技术实现异步块执行,支持高频控制,并在资源受限的机器人平台上实现鲁棒真实部署。

原文摘要 · Abstract (English)

Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily needed to specify task intent rather than to be repeatedly processed during high-frequency low-level execution. Motivated by this separation, we propose a cerebello-thalamic-inspired vision-action model (CT-VAM) for efficient task-conditioned visuomotor control. CT-VAM acts as a compact local execution policy that predicts action chunks from dualview visual observations, proprioception, and a lightweight task condition, potentially enabling a practical cloud-edge paradigm in which high-level semantic reasoning can be handled by large models while fast closed-loop control runs on local hardware. To fuse heterogeneous inputs effectively, CT-VAM introduces TARS (Thalamic Action Routing Stream), a stream-separated conditional attention decoder that independently routes action, visual and task streams, preventing dense sensory tokens from overwhelming compact task-relevant conditions. With only 68M parameters, CT-VAM achieves LIBERO success rates competitive with substantially larger VLA models, while reducing inference latency. Together with flow-consistent inpainting for asynchronous chunk execution, CT-VAM supports high-frequency control and demonstrates robust realworld deployment on resource-constrained robotic platforms.

机器人控制轻量化模型云边协同视觉动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。