arXiv:2606.10180cs.ROcs.AI2026-06

用键盘等简单输入实时控制机器人动作,无需重训即可提升任务成功率。

Flow Control: Steering Vision-Language-Action Models with Simple Real-Time Inputs

论文配图:Flow Control: Steering Vision-Language-Action Models with Simple Real-Time Inputs
图 1 · 摘自论文原文
  • 通过通用输入直接调控视觉语言动作模型的输出行为。
  • 用户输入可精准引导动作,显著提高任务成功率与完成速度。
  • 适合希望快速干预机器人操作的非专业用户使用。

我们提出视觉-语言-动作(VLA)模型的流动控制方法,通过通用输入(如键盘)在实时中简便有效地引导VLA动作,无需重新训练或微调。该方法使粗糙的用户输入能有效对齐用户意图,将输入转化为符合训练阶段专家动作分布的动作样本,从而生成高质量且忠实反映用户意图的动作。实验表明,流动控制具备多项优良特性:(1)能准确、实时地响应用户输入控制机器人动作;(2)对不理想的用户输入具有鲁棒性;(3)显著提升任务成功率并加快完成速度;(4)基于流动控制轨迹微调可进一步优化自主策略。这些结果共同提供了一种简单直观的用户干预方式,显著提升任务表现。

原文摘要 · Abstract (English)

We introduce flow control of vision-language-action (VLA) models, a simple and effective way to steer VLA actions in real-time through generic inputs, such as a keyboard. This method can be used out-of-the-box and does not require retraining or fine-tuning VLAs. It enables relatively crude user inputs to steer a VLA to align with user intent. The VLA transforms these inputs into action samples drawn from the VLA expert action distribution learned during training, so that the generated actions are high quality (conformity to the action expert distribution) and high fidelity (reflecting the user's intent). We demonstrate that flow control has many desirable properties: (1) flow control accurately and responsively steers robot actions with user inputs, (2) it is robust to suboptimal user inputs, (3) it enables users to steer VLAs to achieve significantly higher success rates and faster task completion, and (4) fine-tuning a VLA on flow control trajectories improves the autonomous policy. Together, these results provide a simple and intuitive way for users to help steer VLA actions, increasing task performance.

机器人控制实时干预VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。