无需训练即可个性化图像风格,保持内容一致且高效。
A Training-Free Style-Personalization via SVD-Based Feature Decomposition
- 基于SVD分解特征,提取风格关键成分进行控制。
- 生成图像风格保真度媲美微调模型,推理更快。
- 适合需要快速部署的个性化图像生成场景。
我们提出一种无需训练的风格个性化图像生成框架,通过分尺度自回归模型在推理阶段实现。方法仅需单张参考风格图即可生成风格化图像,同时保持语义一致性并减少内容泄露。通过对生成过程的逐步分析,我们发现内部特征的主奇异值编码了风格相关成分。基于此,引入两个轻量级控制模块:主特征混合(Principal Feature Blending),通过SVD特征重建实现精确风格调节;结构注意力校正(Structural Attention Correction),利用内容引导的注意力修正提升细粒度阶段的结构稳定性。无需额外训练,大量实验表明,该方法在风格保真度和提示保真度上达到与微调基线相当的水平,同时具备更快的推理速度和更强的部署灵活性。
原文摘要 · Abstract (English)
We present a training-free framework for style-personalized image generation that operates during inference using a scale-wise autoregressive model. Our method generates a stylized image guided by a single reference style while preserving semantic consistency and mitigating content leakage. Through a detailed step-wise analysis of the generation process, we identify a pivotal step where the dominant singular values of the internal feature encode style-related components. Building upon this insight, we introduce two lightweight control modules: Principal Feature Blending, which enables precise modulation of style through SVD-based feature reconstruction, and Structural Attention Correction, which stabilizes structural consistency by leveraging content-guided attention correction across fine stages. Without any additional training, extensive experiments demonstrate that our method achieves competitive style fidelity and prompt fidelity compared to fine-tuned baselines, while offering faster inference and greater deployment flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。