arXiv:2604.20705cs.CV2026-04被引 1

用图像自监督生成可验证奖励,提升多模态大模型视觉推理能力。

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

论文配图:SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 将视觉自监督任务转化为可验证的视觉谜题,用于强化学习后训练。
  • 在多个多模态理解与推理基准上显著提升模型性能。
  • 无需人工标注或外部模型,适合大规模视觉增强型模型训练。

基于可验证奖励的强化学习(RLVR)在提升多模态大语言模型(MLLM)推理能力方面展现出巨大潜力。然而,其依赖语言中心先验和昂贵的人工标注,限制了MLLM的内在视觉理解能力及可扩展的奖励设计。本文提出SSL-R1,一种通用的自监督强化学习框架,直接从图像中提取可验证奖励。我们重新审视视觉领域的自监督学习(SSL),将广泛使用的SSL任务重构为一组可验证的视觉谜题,用于强化学习后训练,无需人类或外部模型监督。在这些任务上训练的MLLM在多模态理解与推理基准上表现显著提升,证明了以视觉为中心的自监督任务在MLLM后训练中的潜力。本工作为设计有效的自监督可验证奖励提供了实用经验,助力强化学习的规模化应用。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large language models (MLLMs). However, the reliance on language-centric priors and expensive manual annotations prevents MLLMs' intrinsic visual understanding and scalable reward designs. In this work, we introduce SSL-R1, a generic self-supervised RL framework that derives verifiable rewards directly from images. To this end, we revisit self-supervised learning (SSL) in visual domains and reformulate widely-used SSL tasks into a set of verifiable visual puzzles for RL post-training, requiring neither human nor external model supervision. Training MLLMs on these tasks substantially improves their performance on multimodal understanding and reasoning benchmarks, highlighting the potential of leveraging vision-centric self-supervised tasks for MLLM post-training. We think this work will provide useful experience in devising effective self-supervised verifiable rewards to enable RL at scale. Project page: https://github.com/Jiahao000/SSL-R1.

多模态强化学习自监督视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。