arXiv:2502.19673cs.CV2025-02被引 3

无需微调即可生成任意主体在任意风格中执行指定动作的图像。

SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization

  • 通过新约束增强主体与风格相似性,减少内容和风格泄露。
  • 设计正交时序聚合机制,实现单图+文本的精准条件控制。
  • 可部署于边缘设备,适合移动端个性化图像生成需求。

扩散模型在生成任务中日益流行,尤其适用于主体与风格的个性化组合。尽管现有方法能生成用户指定主体在自定义风格中执行文本引导动作的图像,但通常需微调,难以在移动设备上运行。因此,无需微调的方法如IP-Adapters逐渐受到关注。然而,这些方法在主体与风格组合方面灵活性不足,或依赖ControlNet,或存在内容与风格泄露问题。为此,我们提出SubZero,一种无需微调即可生成任意主体、任意风格、任意动作的框架。我们引入新的约束以增强主体与风格相似性并减少泄露;在去噪模型的交叉注意力块中设计正交时序聚合方案,有效结合文本提示与单张主体及风格图像进行条件控制;还提出定制化内容与风格投影器的训练方法,降低泄露。大量实验证明,该方法在边缘设备上运行良好,显著优于当前最先进的主体、风格与动作组合生成方法。

原文摘要 · Abstract (English)

Diffusion models are increasingly popular for generative tasks, including personalized composition of subjects and styles. While diffusion models can generate user-specified subjects performing text-guided actions in custom styles, they require fine-tuning and are not feasible for personalization on mobile devices. Hence, tuning-free personalization methods such as IP-Adapters have progressively gained traction. However, for the composition of subjects and styles, these works are less flexible due to their reliance on ControlNet, or show content and style leakage artifacts. To tackle these, we present SubZero, a novel framework to generate any subject in any style, performing any action without the need for fine-tuning. We propose a novel set of constraints to enhance subject and style similarity, while reducing leakage. Additionally, we propose an orthogonalized temporal aggregation scheme in the cross-attention blocks of denoising model, effectively conditioning on a text prompt along with single subject and style images. We also propose a novel method to train customized content and style projectors to reduce content and style leakage. Through extensive experiments, we show that our proposed approach, while suitable for running on-edge, shows significant improvements over state-of-the-art works performing subject, style and action composition.

图像生成扩散模型零样本边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。