arXiv:2603.21138cs.CV2026-03

用视觉提示强化学习,让生成模型更懂任务,零样本分类性能提升4.7%。

Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues

  • 引入基于结果奖励的强化学习,让生成模型自进化。
  • 在三个基准上达到新最佳,性能提升4.7%。
  • 适合研究生成式零样本学习与强化学习融合的学者。

生成式零样本学习(ZSL)近年来取得进展,但合成特征常缺乏任务相关性,且仅依赖语义原型难以建模视觉差异大的相似类别。为此,本文提出RLVC框架:通过基于结果的奖励驱动生成模型自演化,提升生成质量;引入类级视觉提示,对齐合成特征与视觉原型并稳定训练过程;设计新型冷启动策略优化训练流程。在三个主流ZSL基准上的实验表明,该方法显著优于现有方法,平均性能提升4.7%,达到当前最优水平。

原文摘要 · Abstract (English)

Recent advances in zero-shot learning (ZSL) have demonstrated the potential of generative models. Typically, generative ZSL synthesizes visual features conditioned on semantic prototypes to model the data distribution of unseen classes, followed by training a classifier on the synthesized data. However, the synthesized features often remain task-agnostic, leading to degraded performance. Moreover, inferring a faithful distribution from semantic prototypes alone is insufficient for classes that are semantically similar but visually distinct. To address these and advance ZSL, we propose RLVC, an outcome-reward reinforcement learning RL framework with visual cues for generative ZSL. At its core, RL empowers the generative model to self-evolve, implicitly enhancing its generation capability. In particular, RLVC updates the generative model using an outcome-based reward, encouraging the synthesis of task-relevant features. Furthermore, we introduce class-wise visual cues that (i) align synthesized features with visual prototypes and (ii) stabilize the RL training updates. For the training process, we present a novel cold-start strategy. Comprehensive experiments and analyses on three prevalent ZSL benchmarks demonstrate that RLVC achieves state-of-the-art results with a 4.7% gain.

零样本学习强化学习生成模型视觉提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。