arXiv:2502.05895cs.CV2025-02被引 3

研究个性化图像生成中采样策略如何提升效果,不依赖微调也能优化生成质量。

Beyond Fine-Tuning: A Systematic Study of Sampling Techniques in Personalized Image Generation

  • 分离采样与微调,系统分析概念轨迹和超类轨迹的影响。
  • 提出决策框架,兼顾生成准确性、保真度与计算效率。
  • 适合关注生成质量与资源平衡的模型开发者和研究人员。

个性化文本到图像生成旨在根据用户定义的概念和描述生成定制化图像。在保持学习概念的保真度与跨场景生成能力之间取得平衡是一项重大挑战。现有方法通常通过不同的微调参数化和改进的采样策略来解决,这些策略在扩散过程中整合超类轨迹。虽然改进的采样策略提供了无需训练、成本较低的解决方案,以增强微调模型的性能,但对这些方法的系统性分析仍然有限。当前方法通常将采样策略与固定微调配置绑定,难以独立评估其对生成结果的影响。为此,我们系统地分析了超越微调的采样策略,探讨概念轨迹和超类轨迹对生成结果的影响。基于此分析,我们提出了一个决策框架,用于评估文本对齐、计算约束和保真度目标,以指导策略选择。该框架可集成多种架构和训练方法,系统优化概念保留、提示遵循性和资源效率。源代码可在 https://github.com/ControlGenAI/PersonGenSampler 获取。

原文摘要 · Abstract (English)

Personalized text-to-image generation aims to create images tailored to user-defined concepts and textual descriptions. Balancing the fidelity of the learned concept with its ability for generation in various contexts presents a significant challenge. Existing methods often address this through diverse fine-tuning parameterizations and improved sampling strategies that integrate superclass trajectories during the diffusion process. While improved sampling offers a cost-effective, training-free solution for enhancing fine-tuned models, systematic analyses of these methods remain limited. Current approaches typically tie sampling strategies with fixed fine-tuning configurations, making it difficult to isolate their impact on generation outcomes. To address this issue, we systematically analyze sampling strategies beyond fine-tuning, exploring the impact of concept and superclass trajectories on the results. Building on this analysis, we propose a decision framework evaluating text alignment, computational constraints, and fidelity objectives to guide strategy selection. It integrates with diverse architectures and training approaches, systematically optimizing concept preservation, prompt adherence, and resource efficiency. The source code can be found at https://github.com/ControlGenAI/PersonGenSampler.

图像生成采样策略个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。