通过干预动作令牌,实时引导机器人完成家庭任务
Steering Autoregressive Vision-Language-Action Policies via Action Token Intervention

- 在推理时直接操作动作令牌空间,无需训练或修改模型
- 任务成功率从10%提升至72.5%,另一项达93.8%
- 适合需要轻量级人机交互的智能机器人应用
我们提出一种名为Token Steering(TS)的方法,通过在自回归视觉-语言-动作(VLA)模型的动作令牌空间中直接干预,动态引导生成的轨迹。该方法将低维用户输入注入模型原生动作令牌表示,使用户能在不修改预训练视觉-语言模型(VLM)架构的前提下影响轨迹生成。由于所有操作均在推理阶段完成,无需额外训练或微调。用户输入仅引导而非覆盖预训练策略,保留了VLA模型所学的灵巧性、平滑性和任务先验。我们在两个家庭操作任务上评估:物体放置后的抽屉关闭和状态感知的物品交换。成功率分别从10.0%提升至72.5%,以及从16.7%提升至93.8%。该接口可为消费级环境中的人机协作提供轻量化、直观的控制方式,提升对行动能力受限人群的可访问性。
原文摘要 · Abstract (English)
We present Token Steering (TS), a method for dynamically steering trajectories generated by an autoregressive vision-language-action (VLA) model through direct intervention in the action-token space. TS injects low-dimensional user inputs into the model's native action-token representation, allowing users to influence trajectory generation without modifying the underlying vision-language model (VLM) architecture. Because TS operates entirely at inference time, it requires no additional training or finetuning. User inputs guide rather than override the pretrained policy, allowing users to influence robot actions while preserving the dexterity, smoothness, and task priors learned by the VLA. We evaluate TS on two household manipulation tasks -- drawer closing after object placement and state-aware object swapping -- and improve success rates from 10.0% to 72.5% and from 16.7% to 93.8%, respectively. By enabling lightweight, intuitive steering over robot foundation models, our interface has the potential to improve human-robot interaction in consumer environments and broaden accessibility for individuals with limited physical control. Project website: https://jasontchan.github.io/token-steering/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。