提出新框架,让视觉语言模型先学上下文再解题,解决数据少时的推理难题。
Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- 分两阶段学习:先理解上下文,再解决问题,避免奖励滥用。
- 在多个基准上超越基线,化学与数学领域提升显著。
- 适合数据稀缺场景,如科研、医学等专业领域的智能模型训练。
近期视觉语言模型(VLMs)通过强化学习(RL)实现卓越推理能力,为连续自进化大型视觉语言模型(LVLMs)提供可行路径。然而,现有方法依赖大量高质量多模态数据,在化学、地球科学及多模态数学等专业领域尤为困难。合成数据与自奖励机制存在分布局限与对齐难题,导致奖励劫持:模型利用高奖励模式,压缩策略熵并破坏训练稳定性。本文提出DoGe(Decouple to Generalize),一种双解耦框架,引导模型优先从上下文而非解题入手,聚焦于合成数据忽略的上下文场景。通过将学习过程解耦为思考者(Thinker)与求解者(Solver)两个模块,合理量化该过程的奖励信号,并采用两阶段强化学习后训练策略:从自由探索上下文过渡到实际任务求解。此外,为提升训练数据多样性,构建了扩展的本域知识语料库和迭代演化的问题种子池。实验表明,该方法在多个基准上持续优于基线,为实现可扩展的自进化LVLM提供有效路径。
原文摘要 · Abstract (English)
Recent vision-language models (VLMs) achieve remarkable reasoning through reinforcement learning (RL), which provides a feasible solution for realizing continuous self-evolving large vision-language models (LVLMs) in the era of experience. However, RL for VLMs requires abundant high-quality multimodal data, especially challenging in specialized domains like chemistry, earth sciences, and multimodal mathematics. Existing strategies such as synthetic data and self-rewarding mechanisms suffer from limited distributions and alignment difficulties, ultimately causing reward hacking: models exploit high-reward patterns, collapsing policy entropy and destabilizing training. We propose DoGe (Decouple to Generalize), a dual-decoupling framework that guides models to first learn from context rather than problem solving by refocusing on the problem context scenarios overlooked by synthetic data methods. By decoupling learning process into dual components (Thinker and Solver), we reasonably quantify the reward signals of this process and propose a two-stage RL post-training approach from freely exploring context to practically solving tasks. Second, to increase the diversity of training data, DoGe constructs an evolving curriculum learning pipeline: an expanded native domain knowledge corpus and an iteratively evolving seed problems pool. Experiments show that our method consistently outperforms the baseline across various benchmarks, providing a scalable pathway for realizing self-evolving LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。