arXiv:2506.03574cs.RO2025-06被引 15

让机器人在执行中实时切换任务,像人一样灵活应对突发指令。

SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models

  • 用执行状态和指令上下文动态调节行为模式
  • 实测任务成功率和交互自然度均优于现有模型
  • 无需额外数据或规划器,适合真实场景应用

部署在动态环境中的机器人不仅需理解多样语言指令,还需在执行过程中灵活响应用户意图的变更。当前视觉-语言-动作(VLA)模型通常假设任务意图固定,无法处理执行中到达的新指令,限制了在零售或家庭等场景中的自然、鲁棒交互。我们提出SwitchVLA,一种统一的、具备执行感知能力的框架,可在不依赖外部规划器或额外切换数据的情况下实现平滑、实时的任务切换。通过将专家示范分割为时间对齐的接触阶段,模型可推断任务进展并相应调整行为。训练一个多行为条件策略,基于不同行为模式生成灵活的动作片段。仿真与真实机器人操作实验表明,SwitchVLA在任务成功率和交互自然性上均超越现有VLA基线,展现出强泛化能力。

原文摘要 · Abstract (English)

Robots deployed in dynamic environments must be able to not only follow diverse language instructions but flexibly adapt when user intent changes mid-execution. While recent Vision-Language-Action (VLA) models have advanced multi-task learning and instruction following, they typically assume static task intent, failing to respond when new instructions arrive during ongoing execution. This limitation hinders natural and robust interaction in dynamic settings, such as retail or household environments, where real-time intent changes are common. We propose SwitchVLA, a unified, execution-aware framework that enables smooth and reactive task switching without external planners or additional switch-specific data. We model task switching as a behavior modulation problem conditioned on execution state and instruction context. Expert demonstrations are segmented into temporally grounded contact phases, allowing the policy to infer task progress and adjust its behavior accordingly. A multi-behavior conditional policy is then trained to generate flexible action chunks under varying behavior modes through conditioned trajectory modeling. Experiments in both simulation and real-world robotic manipulation demonstrate that SwitchVLA enables robust instruction adherence, fluid task switching, and strong generalization-outperforming prior VLA baselines in both task success rate and interaction naturalness.

机器人任务切换多模态动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。