统一模型中联合训练语言与扩散模型,提升多模态理解与生成能力。
UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- 构建统一框架,同时优化语言理解和图像生成能力。
- 定义六种场景,提供统一模型强化学习的系统性基准。
- 适合研究多模态生成与交互的学者参考。
我们提出UniRL-Zero,一种统一的强化学习框架,可同时增强多模态语言模型的理解与推理能力、扩散模型的多媒体生成能力,以及两者间的协同交互能力。本工作定义了六种统一模型强化学习场景,为统一理解与生成模型的强化学习提供了系统性基准。代码已开源:https://github.com/G-U-N/UniRL。
原文摘要 · Abstract (English)
We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。