统一强化学习让视觉语言模型同时提升推理与感知能力
One RL to See Them All: Visual Triple Unified Reinforcement Learning
- 构建三重协同训练框架,解决多模态强化学习中的奖励歧义与稀疏问题
- 在8个任务上联合训练的Orsta模型性能超越专用模型组合
- 适合关注多任务统一训练与视觉语言模型优化的研究者
强化学习正成为后训练视觉语言模型的重要方向,但统一多模态强化学习的公开训练方法仍不成熟,尤其在异构推理与感知密集型任务中。本文提出V-Triune,一种面向统一多模态强化学习的视觉三重统一强化学习方法。其围绕三个协同抽象展开:样本级奖励路由、验证器级结果验证和源级诊断。其中,动态IoU实现定位相关奖励塑造,在宽松阈值下避免奖励歧义,在严格阈值下缓解奖励稀疏性。基于V-Triune,我们开发了Orsta(7B、32B)系列模型,联合训练于八个推理与感知任务。在相同预算下,统一训练表现匹配或优于专用模型混合方案。最终的Orsta模型在MEGA-Bench上优于其基础模型,相较强得多任务强化学习视觉语言模型基线表现良好,并在广泛下游基准上实现性能迁移。结果表明,统一强化学习可在单一视觉语言模型强化学习流程中同时提升推理与感知能力。V-Triune系统及Orsta模型已开源,地址为https://github.com/MiniMax-AI/One-RL-to-See-Them-All。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain much less mature, especially for heterogeneous reasoning and perception-heavy tasks. We propose V-Triune, a Visual Triple Unified Reinforcement Learning methodology for unified multimodal RL. It organizes training around three coordinated abstractions: Sample-Level Reward Routing, Verifier-Level Outcome Verification, and Source-Level Diagnostics. Within this methodology, Dynamic IoU provides localization-specific reward shaping that avoids reward ambiguity under loose thresholds and reward sparsity under strict ones. Built on V-Triune, we develop Orsta (7B, 32B), a family of models jointly trained on eight reasoning and perception tasks. Under matched budgets, unified training matches or outperforms specialist mixtures. The final Orsta models improve over their backbones on MEGA-Bench, compare favorably with strong multi-task RL-VLM baselines, and transfer these gains to a broad set of downstream benchmarks. These results show that unified RL can improve both reasoning and perception within a single VLM RL pipeline.The V-Triune system, along with the Orsta models, is publicly available at https://github.com/MiniMax-AI/One-RL-to-See-Them-All.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。