arXiv:2603.08124cs.ROcs.AI2026-03被引 4

受大脑结构启发,构建可计算感知的视觉-语言-动作系统

SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action

  • 分三模块模拟脑区:皮层保持稳定,脑桥融合感知与意图,小脑快速解码动作
  • 实验显示训练时间缩短40%(7.5→4.5小时),成功率提升至92.5%
  • 模块化设计支持灵活升级,适合机器人控制与高效强化学习场景

我们通过神经科学启发的三元架构重新思考视觉-语言-动作系统。生物学上,皮层提供稳定高层多模态先验并保持冻结;脑桥适配器将皮层特征与实时本体感觉输入融合,并编译意图为执行就绪的标记;小脑(ParaCAT)执行快速并行分类解码以实现在线控制,采用迟滞/指数移动平均/温度/熵机制保障稳定性。固定比例调度与两阶段特征缓存使系统具备计算感知能力且可复现。受主动聚焦视觉启发,腕部感兴趣区域通过校准投影与末端执行器几何关联,提供运动稳定、高分辨率视图,对微小姿态变化敏感,补充主视图全局上下文。该设计模块化:仅更新皮层时只需重训脑桥;更换机器人仅需训练小脑;仅小脑强化学习可优化控制而不影响高层语义。作为概念与协议论文,我们提出在相同条件下(GPU、分辨率、批大小)的定时协议以验证预期效率提升。初步实验表明,在官方N1.5仅头部训练下,分段特征缓存可将训练时间从7.5小时降至4.5小时,平均成功率由86.5%提升至92.5%;SaiVLA0在该设置下达到99.0%均值成功率。

原文摘要 · Abstract (English)

We revisit Vision-Language-Action through a neuroscience-inspired triad. Biologically, the Cerebrum provides stable high-level multimodal priors and remains frozen; the Pons Adapter integrates these cortical features with real-time proprioceptive inputs and compiles intent into execution-ready tokens; and the Cerebellum (ParaCAT) performs fast, parallel categorical decoding for online control, with hysteresis/EMA/temperature/entropy for stability. A fixed-ratio schedule and two-stage feature caching make the system compute-aware and reproducible. Inspired by active, foveated vision, our wrist ROIs are geometrically tied to the end-effector via calibrated projection, providing a movement-stabilized, high-resolution view that is sensitive to fine-grained pose changes and complements the global context of the main view. The design is modular: upgrading the Cerebrum only retrains the Pons; changing robots only trains the Cerebellum; cerebellum-only RL can further refine control without touching high-level semantics. As a concept-and-protocol paper with preliminary evidence, we outline a timing protocol under matched conditions (GPU, resolution, batch) to verify anticipated efficiency gains. We also report preliminary LIBERO evidence showing that split feature caching reduces training time (7.5h to 4.5h) and improves average success (86.5% to 92.5%) under official N1.5 head-only training, and that SaiVLA0 reaches 99.0% mean success.

视觉-语言-动作脑启发模型机器人控制模块化架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。