arXiv:2510.24055cs.ROcs.LG2025-10被引 3

用语言引导视觉表征与专家分工策略,提升机器人多任务操作鲁棒性

Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation

  • 通过语言-视觉联合表征,让机器人理解相似外观任务的差异
  • 在真实机器人上实现79%平均成功率,比基线高21%
  • 适合需要多任务、强泛化能力的机器人控制场景

感知模糊性和任务冲突限制了基于模仿学习的多任务机器人操作。我们提出一个结合语言条件视觉表征(LCVR)模块和语言条件混合专家密度策略(LMoE-DP)的框架。LCVR通过将视觉特征与语言指令对齐,解决感知模糊问题,使系统能区分视觉相似的任务。为缓解任务冲突,LMoE-DP采用稀疏专家结构,分别学习不同模态的动作分布,并通过梯度调制稳定训练。在真实机器人基准测试中,LCVR使基于Transformer的动作分块(ACT)和扩散策略(DP)的成功率分别提升33.75%和25%。完整框架实现79%的平均成功率,优于先进基线21%。结果表明,语义对齐与专家专业化结合可实现高效、鲁棒的多任务操作。

原文摘要 · Abstract (English)

Perceptual ambiguity and task conflict limit multitask robotic manipulation via imitation learning. We propose a framework combining a Language-Conditioned Visual Representation (LCVR) module and a Language-conditioned Mixture-ofExperts Density Policy (LMoE-DP). LCVR resolves perceptual ambiguities by grounding visual features with language instructions, enabling differentiation between visually similar tasks. To mitigate task conflict, LMoE-DP uses a sparse expert architecture to specialize in distinct, multimodal action distributions, stabilized by gradient modulation. On real-robot benchmarks, LCVR boosts Action Chunking with Transformers (ACT) and Diffusion Policy (DP) success rates by 33.75% and 25%, respectively. The full framework achieves a 79% average success, outperforming the advanced baseline by 21%. Our work shows that combining semantic grounding and expert specialization enables robust, efficient multi-task manipulation

机器人操作多任务学习语言引导专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。