通过社区LoRA挖掘实现风格与内容分离的可控图像生成。
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

- 利用社区LoRA作为风格和内容的组合锚点,构建大规模三元组数据集。
- 提出两阶段课程学习机制,有效抑制风格泄露,提升内容保真度。
- 适用于需要精准控制风格与内容的图像生成任务,如艺术创作与设计辅助。
风格-内容双参考生成旨在合成一张保持内容参考结构与语义,同时采用独立风格参考风格的图像。尽管近期取得进展,该任务仍具挑战性,因模型需平衡内容保真、风格对齐与指令遵循,避免风格参考中的语义泄露。主要瓶颈在于缺乏大规模、具有清晰风格-内容分离且覆盖广泛长尾风格的三元组数据。本文提出FreeStyle,一种基于社区LoRA挖掘的可扩展双参考生成框架。我们将社区LoRA视为风格与内容的组合锚点,设计严谨的生成与过滤流程,在多个基础模型上构建大规模风格参考与内容参考三元组。为解决内容泄露问题,采用两阶段课程学习,包含阶段特异性解耦机制:注意力级增强约束抑制风格迁移阶段的风格参考泄露;频域感知的RoPE调制策略针对更难的双参考阶段中的位置对应泄露。我们还引入一个基准测试,涵盖风格参考与双参考生成,评估风格相似性、内容保真、美学质量、指令遵循及泄露抑制。该基准包含风格不变的内容对齐得分(CAS),并引入校准的基于VLM的拒绝得分以评估生成可靠性与泄露抑制。大量实验表明,所提模型在风格对齐、内容保真与泄露抑制之间达到良好平衡。
原文摘要 · Abstract (English)
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style reference.Despite recent progress, this setting remains challenging because models must balance content fidelity, style alignment, and instruction following avoiding semantic leakage from the style reference.A key bottleneck is the lack of large-scale triplet data with clean content-style separation and broad long-tail style coverage.In this work, we propose FreeStyle, a scalable dual-reference generation framework based on community LoRA mining.We treat community LoRAs as compositional anchors for style and content, and design a rigorous generation and filtering pipeline to construct large-scale Style-Reference and Content-Reference triplets across multiple base models.To address content leakage, we adopt a two-stage curriculum with stage-specific disentanglement mechanisms: an attention-level enrichment constraint that suppresses style-reference leakage in the style-transfer stage, and a frequency-aware RoPE modulation strategy that targets positional-correspondence-based leakage in the harder dual-reference stage.We also introduce a benchmark covering both style-reference and dual-reference generation, with evaluations on style similarity, content preservation, aesthetics, instruction following, and leakage rejection. The benchmark incorporates a style-invariant Content Alignment Score (CAS) and introduces a calibrated VLM-based Rejection Score for evaluating generation reliability and leakage suppression.Extensive experiments show that our model achieves a strong balance among style alignment, content preservation, and leakage suppression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。