arXiv:2605.03403cs.CVcs.LG2026-05

用强化学习提升视觉语言模型测试时适应能力

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning

论文配图:GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
图 1 · 摘自论文原文
  • 将分组相对策略优化引入测试时调优,无需真实标签
  • 在分布偏移下性能优于现有方法,提升显著
  • 适合需要强泛化能力的视觉语言模型应用

组相对策略优化(GRPO)在大语言模型和视觉语言模型的后训练中表现优异。本文提出针对测试时适应(TTA)的GRPO-TTA方法,将类别特定提示预测重构为分组策略优化问题。通过从CLIP相似度分布中采样前K个类别候选,构建输出组,实现无真实标签的概率驱动优化。设计了对齐奖励与分散奖励,引导视觉编码器有效调优。在多个基准上的实验证明,GRPO-TTA持续优于现有测试时适应方法,尤其在自然分布偏移下性能提升更明显。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models. It raises a question of whether the GRPO also significantly promotes the test-time adaptation (TTA) of vision language models. In this paper, we propose Group Relative Policy Optimization for Test-Time Adaptation (GRPO-TTA), which adapts GRPO to the TTA setting by reformulating class-specific prompt prediction as a group-wise policy optimization problem. Specifically, we construct output groups by sampling top-K class candidates from CLIP similarity distributions, enabling probability-driven optimization without access to ground-truth labels. Moreover, we design reward functions tailored to test-time adaptation, including alignment rewards and dispersion rewards, to guide effective visual encoder tuning. Extensive experiments across diverse benchmarks demonstrate that GRPO-TTA consistently outperforms existing test-time adaptation methods, with notably larger performance gains under natural distribution shifts.

视觉语言模型测试时调优强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。