用强化学习让多模态生成更智能,支持输入输出全多模态。
M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation
- 用强化学习训练图像插入模块,实现可控语义对齐。
- 30亿参数轻量模型延迟低,质量超越基线。
- 适合需要多模态输出的智能创作与交互系统。
当前多模态检索增强生成(MRAG)研究虽支持多样多模态输入,但仅限单模态输出,限制了表达能力和实际应用。真实场景常需多模态输入与输出以实现有效沟通和基于事实的推理。受大语言模型复杂推理任务中强化学习(RL)成功启发,我们采用RL作为解决多步、目标驱动型多模态输出生成挑战的系统性方法。本文提出M2IO-R1框架,支持多模态输入与输出的多模态检索增强多模态生成(MRAMG)。核心是基于强化学习的图像插入模块Inserter-R1-3B,通过分组相对策略优化训练,实现图像选择与放置的可控且语义对齐。实验表明,该轻量级30亿参数插入器在显著降低延迟的同时,展现出强推理能力,质量与效率均优于基线。
原文摘要 · Abstract (English)
Current research on Multimodal Retrieval-Augmented Generation (MRAG) enables diverse multimodal inputs but remains limited to single-modality outputs, restricting expressive capacity and practical utility. In contrast, real-world applications often demand both multimodal inputs and multimodal outputs for effective communication and grounded reasoning. Motivated by the recent success of Reinforcement Learning (RL) in complex reasoning tasks for Large Language Models (LLMs), we adopt RL as a principled and effective paradigm to address the multi-step, outcome-driven challenges inherent in multimodal output generation. Here, we introduce M2IO-R1, a novel framework for Multimodal Retrieval-Augmented Multimodal Generation (MRAMG) that supports both multimodal inputs and outputs. Central to our framework is an RL-based inserter, Inserter-R1-3B, trained with Group Relative Policy Optimization to guide image selection and placement in a controllable and semantically aligned manner. Empirical results show that our lightweight 3B inserter achieves strong reasoning capabilities with significantly reduced latency, outperforming baselines in both quality and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。