arXiv:2603.13335cs.CVcs.AI2026-03被引 2

提出Info-VLA框架,缓解机器人持续学习中的遗忘问题。

Information-Theoretic Constraints for Continual Vision-Language-Action Alignment

  • 用教师模型构建稳定对齐锚点,保持跨模态结构
  • 通过互信息最大化维持视觉与语言依赖关系
  • 适合需要长期学习新技能的机器人系统

在开放式的机器人环境中,视觉-语言-动作(VLA)模型需持续学习新技能,但会遭遇严重灾难性遗忘。我们观察到,这种性能退化与跨模态信息结构的恶化相关:视觉观测、语言指令与动作之间的依赖关系在持续适应过程中逐渐扩散。然而,现有持续学习方法无法保留此类跨模态依赖。为此,我们提出Info-VLA,一种基于信息保持的持续学习框架,通过两个互补约束维持跨模态信息结构。回放锚点对比学习从冻结的教师模型中构建稳定的对齐锚点,保留表示空间中的跨模态对齐。跨模态互信息最大化通过互信息约束进一步保持视觉与语言表示间的依赖结构。联合维护历史对齐与跨模态依赖信息,Info-VLA 在持续学习中实现了稳定性与可塑性的平衡。在LIBERO数据集上的实验表明,Info-VLA 在任务保留和适应能力方面显著优于现有方法。

原文摘要 · Abstract (English)

When deployed in open-ended robotic environments, Vision--Language--Action (VLA) models need to continually acquire new skills, yet suffer from severe catastrophic forgetting. We observe that this degradation is related to the deterioration of cross-modal information structure, where dependencies among visual observations, language instructions, and actions progressively diffuse during continual adaptation. But existing continual learning methods fail to preserve such cross-modal information dependencies. Thus, we propose Info-VLA, an information-preserving continual learning framework that maintains cross-modal information structure through two complementary constraints. Replay Anchor Contrastive Learning constructs stable alignment anchors from a frozen teacher model, preserving cross-modal alignment in the representation space. Cross-Modal Mutual Information Maximization further preserves dependency structure between visual and language representations through mutual information constraints. By jointly preserving historical alignment and cross-modal dependency information, Info-VLA balances stability and plasticity during continual learning. Furthermore, experiments on the LIBERO show that Info-VLA significantly outperforms existing methods in both task retention and adaptation.

持续学习跨模态对齐机器人信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。