arXiv:2607.07173cs.CV2026-07

针对个性化图像生成,提出分阶段适配与分布校准,提升身份一致性与多样性。

Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation

论文配图:Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation
图 1 · 摘自论文原文
  • 分阶段低秩适配,按去噪阶段动态调整参数更新强度。
  • 推理时校准候选样本分布,提升文本对齐与视觉多样性。
  • 适合需要高保真个性化生成的研究者与应用开发者。

基于扩散模型的主体驱动个性化文生图生成需从少量参考图像中学习特定主体,保持身份一致、遵循新文本提示并维持样本多样性。现有优化方法通过全微调、文本嵌入优化或低秩参数更新实现个性化;PaRa进一步从参数秩压缩角度约束个性化。然而,统一的低秩约束或固定适配强度无法区分不同去噪阶段的能力需求。此外,仅依赖身份相似度的推理阶段候选选择会压缩视觉表示空间中的样本分布。本文将问题分解为两个互补组件:SPaRa(训练侧分阶段低秩适配)与DCAL(推理侧分布校准候选选择),构成SPaRa-DCAL联合框架。理论分析表明,时间步依赖缩放控制低秩适配器的有效扰动幅度,而身份偏差候选选择在特定条件下限制了特征围绕参考中心的半径。在SDXL和DreamBooth 30主体协议下的可审计实验显示,DCAL在固定LoRA候选池上提升了1-LPIPS、CLIP-I、DINO-I与CLIP-T指标,但与CLIP/DINO成对多样性及成对LPIPS存在明显权衡。结果表明,个性化生成应综合评估身份一致性、文本对齐与表征多样性,而非仅依赖身份指标。

原文摘要 · Abstract (English)

Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and maintaining sample diversity. Existing optimization-based methods instantiate subject adaptation through full fine-tuning, textual embedding optimization, or low-rank parameter updates; PaRa further constrains personalization from the perspective of parameter rank reduction. However, a uniform low-rank constraint or a uniform adapter strength cannot explicitly distinguish the capacity requirements of different denoising stages. Moreover, inference-time candidate selection driven mainly by identity similarity may compress the selected samples in the visual representation space. We decompose the problem into two complementary components: SPaRa denotes training-side stage-aware low-rank adaptation, DCAL denotes inference-side distribution-calibrated candidate selection, and SPaRa-DCAL denotes the combined framework. Theoretical analysis shows that timestep-dependent scaling controls the effective perturbation magnitude of a low-rank adapter, while identity-biased candidate selection restricts the radius of selected features around the reference center under explicit conditions. Auditable experiments under the SDXL and DreamBooth 30-subject protocol show that DCAL improves 1-LPIPS, CLIP-I, DINO-I, and CLIP-T on a fixed LoRA candidate pool, while revealing a clear trade-off with CLIP/DINO pairwise diversity and pairwise LPIPS. These results indicate that personalized generation should be evaluated through identity consistency, text alignment, and representation diversity rather than identity metrics alone.

文生图个性化生成扩散模型低秩适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。