用强化学习提升视觉语言模型的通用推理能力。
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- 自动生成带推理路径的图文问答对,支持多领域任务。
- 构建12万+样本的WeThink数据集,覆盖18个来源与多种问题类型。
- 结合规则验证与模型评估的混合奖励机制,提升训练效率。
基于文本推理模型DeepSeek-R1的成功经验,将此类能力拓展至多模态推理具有巨大潜力。尽管已有研究尝试将类似DeepSeek-R1的强化学习(RL)训练范式应用于多模态大语言模型(MLLM),但主要聚焦于数学和视觉感知等特定任务,一个关键问题仍待解决:如何通过强化学习实现通用视觉语言推理?为此,我们提出三项核心贡献:(1) 一种可扩展的多模态问答生成流水线,能从图像中自动合成上下文相关、以推理为核心的问答对;(2) 开源的WeThink数据集,包含超过12万组多模态问答对,并附有标注的推理路径,数据源自18个不同来源,覆盖多种问题领域;(3) 在该数据集上对强化学习进行系统探索,采用结合规则验证与模型评估的混合奖励机制,显著提升跨任务域的训练效率。在14个多样化的MLLM基准测试中,WeThink数据集显著提升了模型性能,涵盖数学推理到各类通用多模态任务。此外,我们证明该自动化数据流水线可持续增加数据多样性,进一步优化模型表现。
原文摘要 · Abstract (English)
Building on the success of text-based reasoning models like DeepSeek-R1, extending these capabilities to multimodal reasoning holds great promise. While recent works have attempted to adapt DeepSeek-R1-style reinforcement learning (RL) training paradigms to multimodal large language models (MLLM), focusing on domain-specific tasks like math and visual perception, a critical question remains: How can we achieve the general-purpose visual-language reasoning through RL? To address this challenge, we make three key efforts: (1) A novel Scalable Multimodal QA Synthesis pipeline that autonomously generates context-aware, reasoning-centric question-answer (QA) pairs directly from the given images. (2) The open-source WeThink dataset containing over 120K multimodal QA pairs with annotated reasoning paths, curated from 18 diverse dataset sources and covering various question domains. (3) A comprehensive exploration of RL on our dataset, incorporating a hybrid reward mechanism that combines rule-based verification with model-based assessment to optimize RL training efficiency across various task domains. Across 14 diverse MLLM benchmarks, we demonstrate that our WeThink dataset significantly enhances performance, from mathematical reasoning to diverse general multimodal tasks. Moreover, we show that our automated data pipeline can continuously increase data diversity to further improve model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。