arXiv:2601.06122cs.CVcs.AI2026-01AAAI被引 1

让视觉语言模型和强化学习互相提升,提高复杂任务的样本效率。

COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

  • 用强化学习生成的数据微调视觉语言模型,增强语义理解。
  • 在多个视觉控制任务中性能显著优于基线方法。
  • 适合研究视觉强化学习与多模态模型协同优化的学者。

视觉强化学习因复杂任务中的高维观测导致样本效率低下。现有工作虽证明视觉语言模型(VLM)可辅助强化学习,但多聚焦于将VLM知识蒸馏给强化学习,忽视了强化学习生成的交互数据对VLM的反向增益。为此,我们提出COVR框架,实现VLM与强化学习策略的协同优化:首先利用强化学习生成的数据微调VLM,以增强其与目标任务一致的语义推理能力;随后,借助优化后的VLM提供动作先验,指导策略学习。为提升微调效率,引入两个关键模块:(1) 探索驱动的动态筛选模块,基于探索程度自适应保留高价值样本;(2) 回报感知的自适应损失权重模块,通过强化学习回报信号量化采样动作不一致性,提升训练稳定性。此外,设计渐进式微调策略以降低资源消耗。大量实验表明,COVR在多个挑战性视觉控制任务中均表现出色。

原文摘要 · Abstract (English)

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language models (VLMs) can assist RL, they often focus on knowledge distillation from the VLM to RL, overlooking the potential of RL-generated interaction data to enhance the VLM. To address this, we propose COVR, a collaborative optimization framework that enables the mutual enhancement of the VLM and RL policies. Specifically, COVR fine-tunes the VLM with RL-generated data to enhance the semantic reasoning ability consistent with the target task, and uses the enhanced VLM to further guide policy learning via action priors. To improve fine-tuning efficiency, we introduce two key modules: (1) an Exploration-Driven Dynamic Filter module that preserves valuable exploration samples using adaptive thresholds based on the degree of exploration, and (2) a Return-Aware Adaptive Loss Weight module that improves the stability of training by quantifying the inconsistency of sampling actions via return signals of RL. We further design a progressive fine-tuning strategy to reduce resource consumption. Extensive experiments show that COVR achieves strong performance across various challenging visual control tasks.

视觉强化学习多模态协同高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。