用智能路由筛选最佳专家,提升机器人团队应对新场景能力
RouterVLA: Budgeted Commissioning and Expert Onboarding for Growing VLA Pools

- 设计路由机制,基于失败反馈动态选择最优专家
- 在预算不变下达成60.53%的外部成功率,比语义筛选高1.64个百分点
- 能自动发现覆盖现有盲区且综合能力强的专家,适合持续扩展的智能系统
机器人团队通常维护多个视觉-语言-动作策略,但仅部署单一全局最优策略。本文研究两个关键决策:为新场景选择哪个专家部署,以及向策略池中新增哪个候选。RouterVLA结合分治清洗先验与结果不相交探测,采用仅对当前策略池无法处理的失败进行奖励的上岗机制。在严格成本匹配的探测预算下,其在未见场景上的成功率达60.53%,较语义短名单方法提升1.64个百分点。两项指标独立收敛至同一组五名专家,表明最能填补现有策略池盲区的候选,也具备广泛适应能力。
原文摘要 · Abstract (English)
Robotic teams often maintain several vision-language-action policies but still deploy one global winner. We study two recurring decisions: which expert to deploy for a new condition and which candidate to add to the pool. RouterVLA combines a split-clean prior and outcome-disjoint probes with onboarding that credits only failures the incumbent pool cannot handle. Under an exactly cost-matched probe budget, it reaches 60.53\% held-out success, a $+1.64\pp$ gain over a semantic shortlist. Both criteria independently converge on the same five experts, confirming that the candidates best positioned to cover the base pool's blind spots are also broadly capable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。