arXiv:2606.27144cs.RO2026-06

让机器人分阶段执行任务更可靠,通过专家分工提升动作准确性

PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies

论文配图:PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 用多个专家分工代替单一动作模型,按任务阶段分配不同专家
  • 在仿真任务中使成功率最高提升9.2%,显著优于现有方法
  • 适合需要多阶段精准控制的机器人操作场景

多阶段机器人操作中的可靠动作生成仍是视觉-语言-动作(VLA)模型的挑战。尽管现有流匹配VLA策略具备强多模态对齐与泛化能力,但通常采用单一共享动作专家,难以捕捉不同执行阶段的特定控制模式。本文提出一种即插即用的相位感知专家混合模块(PAMAE),作为迈向更可靠阶段一致动作生成的重要一步。PAMAE用稀疏专家混合替代原流匹配动作专家,同时保留预训练VLA主干。其引入相位感知路由机制,利用执行阶段线索分配动作生成任务,辅以轻量级相位预测头和路由对齐目标。为稳定专家分工,采用两阶段训练:先在标准流匹配损失下预热专家模块,再在辅助监督下优化阶段一致性路由。在多阶段操作仿真任务中,PAMAE相比强基线最高提升任务成功率9.2%。消融实验表明,相位监督路由与分阶段优化均对性能提升至关重要。结果表明,阶段一致的专家分配是提升流匹配VLA策略可靠性与动作质量的有效机制。

原文摘要 · Abstract (English)

Reliable action generation for multi-stage robotic manipulation remains challenging for Vision-Language-Action (VLA) models. While existing flow-matching VLA policies offer strong multimodal grounding and generalization, they typically employ a single shared action expert, limiting their ability to capture phase-specific control patterns across distinct execution stages. We propose a plug-and-play Phase-Aware Mixture-of-Experts Action Module (PAMAE), as a step towards more reliable phase-consistent action generation. PAMAE replaces the original flow-matching action expert with a sparse expert mixture while preserving the pretrained VLA backbone. PAMAE introduces a phase-aware router that leverages execution-phase cues to allocate action generation across experts, supported by a lightweight phase prediction head and a routing alignment objective. To stabilize specialization, we adopt a two-stage training scheme that first warms up the expert module under the standard flow-matching loss and then optimizes phase-consistent routing under auxiliary supervision. On multi-stage manipulation simulation tasks, PAMAE improves task success by up to \textbf{9.2\%} over strong VLA baselines. Further ablations show that both phase-supervised routing and staged optimization are essential for the observed gains. Our results highlight phase-consistent expert allocation as an effective mechanism for improving the reliability and action quality of flow-matching VLA policies.

机器人控制动作生成专家混合多阶段任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。