通过动态分组残差优化,让视觉语言动作模型跨任务通用性更强。
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization

- 基于信息论原理捕捉跨任务潜在表示
- 多强化学习残差动态调整策略,提升泛化能力
- 适合需要多任务适应的机器人控制场景
近期强化学习进展为视觉-语言-动作(VLA)模型优化提供了理论基础,推动其从轨迹模仿转向环境中的主动学习。尽管控制精度提升,多数强化学习优化器仍局限于特定任务,使VLA模型退化为窄范围任务的过拟合策略。本文深入分析该现象,强调跨任务特征表示对提升模型泛化能力的重要性。为此,提出DyGRO-VLA两阶段优化框架:1)基于信息论原则有效捕获跨任务潜在表示;2)通过强化学习残差混合机制动态优化策略。该方法使优化器能利用任务相关潜信息,同时在训练过程中战略性缓解表示学习的干扰。在LIBERO、RoboTwin2基准上评估,并进一步在真实世界中验证,结果显示在多任务训练和分布外测试下均显著优于强基线。
原文摘要 · Abstract (English)
Recent progress in Reinforcement Learning (RL) provides a principled approach to optimizing Vision-Language-Action (VLA) models, facilitating a shift from trajectory imitation to active learning in the task environment. Despite improvements in control precision, most RL optimizers remain task-specific, which reduces VLA models from generalist controllers to policies that overfit to a narrow set of tasks. In this study, we conduct an in-depth analysis of this phenomenon and highlight the importance of cross-task feature representations for improving the generalizability of VLA models. Motivated by this finding, we introduce DyGRO-VLA, a two-stage optimization framework that 1) effectively captures cross-task latent representations based on information-theoretic principles, and 2) dynamically refines policy optimization via a mixture-of-RL-residuals. DyGRO-VLA enables the RL optimizer to exploit task-relevant latent information while strategically mitigating adverse interference on the learned representations throughout the optimization process. We evaluate our approach on LIBERO, RoboTwin2 benchmarks, and further validate it on real world, demonstrating consistent improvements over strong baselines under multi-task training and distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。