arXiv:2511.08061cs.CVcs.AI2025-11被引 1

解决扩散模型生成中身份一致与提示多样性冲突问题。

Taming Identity Consistency and Prompt Diversity in Diffusion Models via Latent Concatenation and Masked Conditional Flow Matching

  • 通过潜在空间拼接与掩码条件流匹配,实现身份保真无需改架构。
  • 两阶段数据蒸馏框架支持大规模高效微调,提升多主体泛化能力。
  • 提出细粒度评估框架CHARIS,从五维度量化生成质量与多样性。

基于主体的图像生成旨在合成同一主体在多样场景下的新图像,同时保持其核心身份特征。实现强身份一致性与高提示多样性存在根本性权衡。本文提出一种基于LoRA微调的扩散模型,采用潜在空间拼接策略联合处理参考图与目标图,并结合掩码条件流匹配(masked Conditional Flow Matching)目标,可在不修改架构的前提下实现鲁棒的身份保留。为支持大规模训练,引入两阶段数据蒸馏框架:第一阶段通过数据恢复与VLM-based过滤,从多样化来源构建紧凑高质量种子数据集;第二阶段利用该数据集进行参数高效微调,从而扩展生成能力至多种主体与场景。最后,针对过滤与质量评估,提出CHARIS框架,从身份一致性、提示遵循性、区域色彩保真度、视觉质量与变换多样性五个关键维度进行细粒度比较。

原文摘要 · Abstract (English)

Subject-driven image generation aims to synthesize novel depictions of a specific subject across diverse contexts while preserving its core identity features. Achieving both strong identity consistency and high prompt diversity presents a fundamental trade-off. We propose a LoRA fine-tuned diffusion model employing a latent concatenation strategy, which jointly processes reference and target images, combined with a masked Conditional Flow Matching (CFM) objective. This approach enables robust identity preservation without architectural modifications. To facilitate large-scale training, we introduce a two-stage Distilled Data Curation Framework: the first stage leverages data restoration and VLM-based filtering to create a compact, high-quality seed dataset from diverse sources; the second stage utilizes these curated examples for parameter-efficient fine-tuning, thus scaling the generation capability across various subjects and contexts. Finally, for filtering and quality assessment, we present CHARIS, a fine-grained evaluation framework that performs attribute-level comparisons along five key axes: identity consistency, prompt adherence, region-wise color fidelity, visual quality, and transformation diversity.

扩散模型图像生成身份保真提示多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。