用强化学习显式对齐图文表示,提升多模态理解与生成能力。
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
- 基于GRPO的强化学习框架,直接优化模型行为实现图文对齐。
- 在多个评测上超越现有统一预训练基线,尤其在细粒度推理任务中表现优异。
- 无需额外编码器或手工设计目标,适合需要可控生成的场景。
统一多模态预训练已成为在单一基础模型中联合建模语言与视觉的有前景范式。然而,现有方法主要依赖隐式或间接对齐信号,在同时支持多模态理解与生成方面仍不理想,尤其在需要细粒度语言-视觉推理和可控生成的场景中表现不足。本文提出LVRPO,一种基于强化学习的图文对齐框架,采用分组相对策略优化(GRPO)显式对齐语言与视觉表征。不同于在表示层引入额外对齐损失,LVRPO通过偏好驱动的强化信号直接优化多模态模型行为,促进语言与视觉在理解与生成任务中的语义一致交互。该方法无需辅助编码器或人工设计的跨模态目标,天然可扩展至多样化多模态能力。实验表明,LVRPO在涵盖多模态理解、生成与推理的广泛基准测试中持续优于强基线模型。
原文摘要 · Abstract (English)
Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely rely on implicit or indirect alignment signals and remain suboptimal for simultaneously supporting multimodal understanding and generation, particularly in settings that require fine-grained language-visual reasoning and controllable generation. In this work, we propose LVRPO, a language-visual reinforcement-based preference optimization framework that explicitly aligns language and visual representations using Group Relative Policy Optimization (GRPO). Instead of introducing additional alignment losses at the representation level, LVRPO directly optimizes multimodal model behaviors through preference-driven reinforcement signals, encouraging consistent and semantically grounded interactions between language and vision across both understanding and generation tasks. This formulation enables effective alignment without requiring auxiliary encoders or handcrafted cross-modal objectives, and naturally extends to diverse multimodal capabilities. Empirically, LVRPO consistently outperforms strong unified-pretraining baselines on a broad suite of benchmarks spanning multimodal understanding, generation, and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。