arXiv:2606.23950cs.CV2026-06中稿 · ECCV

解决主体生成中身份一致与多样性矛盾,提升图像多样性。

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

论文配图:DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
图 1 · 摘自论文原文
  • 用解耦特征分离结构多样性和身份一致性
  • 提出负自相似度衡量结构多样性,提升生成多样性
  • 适合需要高多样性主体生成的场景

主体驱动图像生成面临“身份-多样性悖论”,即强身份保持常导致输出僵化、多样性低。我们提出一种后训练框架DivRL,通过利用鲁棒相似性模型的解耦视觉特征,同时优化身份一致性和结构多样性。具体地,引入负自相似度度量(nSSM)量化结构多样性,视觉语义匹配(VSM)评估身份一致性。提出“探索-抑制”策略,将VSM作为门控约束:模型自由探索结构多样的配置,仅违反身份阈值的样本通过二次铰链损失惩罚。这将身份保持从竞争目标转化为可行性约束,使nSSM与VSM可协同优化。实验表明,该方法有效推动模型生成既一致又多样的图像,在保持相近身份一致性的同时显著提升结构多样性。

原文摘要 · Abstract (English)

Subject-driven image generation faces an "Identity-Diversity Paradox", where strong identity preservation often leads to rigid and low-diversity outputs. We propose a post-training framework called DivRL that jointly optimizes identity consistency and structural diversity simultaneously by leveraging disentangled visual features from a robust similarity model. Specifically, we introduce a Negative Self-Similarity Measure (nSSM) to quantify structural diversity, and Visual Semantic Matching (VSM) to evaluate identity consistency. We propose an "Explore-and-Suppress" strategy that treats VSM as a gated constraint: the model freely explores structurally diverse configurations, and only samples that violate the identity threshold are penalized via a quadratic hinge loss. This converts identity preservation from a competing objective into a feasibility constraint, allowing nSSM and VSM to improve jointly. Experiments demonstrate that our method effectively pushes the model to generate both consistent and diverse images and improves structural diversity while maintaining comparable identity consistency through a gated optimization formulation.

图像生成多样性身份保持强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。