arXiv:2508.11433cs.CV2025-08AAAI被引 10

让统一多模态大模型零样本生成个性化图像

MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation

  • 用跨模态思维链解析用户图文,实现主体概念定位
  • 零样本下生成图像主体一致且对齐文本描述
  • 适合需要快速个性化图像生成的场景

统一架构的多模态大语言模型在视觉-语言任务中表现优异,但其在个性化图像生成中的应用仍面临挑战。现有方法通常依赖特定主体,需为每个新主体进行数据密集型微调,难以扩展。本文提出MM-R1框架,通过跨模态思维链(X-CoT)推理策略,将个性化生成视为整合的视觉推理与生成过程:(1) 通过解析用户提供的图像和上下文线索,定位主体概念;(2) 基于提取的主体表征和用户提示生成个性化图像。为进一步增强推理能力,采用分组奖励近端策略优化(GRPO)显式对齐生成过程。实验表明,MM-R1在零样本条件下即可释放统一多模态大模型的个性化生成潜力,生成图像具备高主体保真度和强文本对齐性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization (GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner.

个性化生成多模态大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。