用强化学习提升多模态模型个性化图文生成能力
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- 基于强化学习替代传统监督微调,优化个性化生成
- 在多概念图像描述任务中显著优于现有方法
- 适合需要高个性化图文生成的场景研究者
近期多模态大语言模型在生成个性化图像描述时仍表现不佳,即使使用高质量标注数据训练。本文发现,现有基于后训练的个性化方法仍存在局限:尽管通过监督微调(SFT)使用大规模标注数据进行调优,但在真实场景如多概念图像描述中仍难以生成准确描述。然而,获取此类复杂场景下的大规模高质量标注数据成本高昂且困难。为解决SFT对数据的依赖问题,我们提出首个基于强化学习(RL)的多模态大模型后训练框架——RePIC。该方法显著提升模型的视觉识别与个性化生成能力,在多概念图像描述任务上持续优于现有SFT基线,有效缓解数据瓶颈。
原文摘要 · Abstract (English)
Recent multi-modal large language models (MLLMs) often struggle to generate personalized image captions, even when trained on high-quality captions. In this work, we observe that such limitations persist in existing post-training-based MLLM personalization methods. Specifically, despite being post-tuned with large-scale caption data through supervised fine-tuning (SFT), these models frequently fail to produce faithful descriptions in real-world scenarios, such as multi-concept image captioning. However, acquiring large-scale, high-quality captions for such complex settings is both costly and difficult. To address the data-centric nature of SFT, we propose a reinforcement learning (RL)-based post-training framework. To the best of our knowledge, this is the first RL-based approach to post-train MLLMs for personalized image captioning. Our method significantly enhances both visual recognition and personalized generation capabilities of MLLMs, and consistently outperforms existing SFT-based baselines, especially in the challenging multi-concept image captioning task. Project page: https://github.com/oyt9306/RePIC
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。