用自监督学习任务生成可靠奖励,提升视觉语言模型推理能力
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- 将图像旋转、掩码重建等自监督任务转为密集奖励信号
- 在多个视觉语言基准上显著提升性能,最高提升12.7%
- 无需人工标注,适用于多模态与图学习场景
视觉语言模型(VLM)虽能融合语言与视觉输入,但常依赖语言先验或使用文本捷径,难以有效利用视觉证据。强化学习(RL)可对齐模型行为,但缺乏可扩展且可靠的奖励机制限制了其应用。本文提出SSL4RL框架,将自监督学习(SSL)任务如图像旋转预测、掩码图像重建转化为可验证的密集奖励信号,避免依赖人类偏好数据或不可靠的AI评估器。实验表明,该方法在视觉中心与跨模态推理任务上均有显著提升,系统消融分析揭示任务难度、模型规模及语义对齐度是关键影响因素,为未来设计提供新原则。此外,该框架在图学习任务中也表现优异,证明其通用性。SSL4RL建立了一种基于可验证自监督目标的多模态对齐新范式。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often fail to utilize visual evidence adequately, either depending on linguistic priors in vision-centric tasks or resorting to textual shortcuts during reasoning. Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs has been hindered by the lack of scalable and reliable reward mechanisms. To overcome this challenge, we propose SSL4RL, a novel framework that leverages self-supervised learning (SSL) tasks as a source of verifiable rewards for RL-based fine-tuning. Our approach reformulates SSL objectives-such as predicting image rotation or reconstructing masked patches-into dense, automatic reward signals, eliminating the need for human preference data or unreliable AI evaluators. Experiments show that SSL4RL substantially improves performance on both vision-centric and vision-language reasoning benchmarks. Furthermore, through systematic ablations, we identify key factors-such as task difficulty, model scale, and semantic alignment with the target domain-that influence the effectiveness of SSL4RL tasks, offering new design principles for future work. We also demonstrate the framework's generality by applying it to graph learning, where it yields significant gains. SSL4RL establishes a versatile and effective paradigm for aligning multimodal models using verifiable, self-supervised objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。