无需微调,通过球面插值实现多风格无缝融合。
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation
- 用球面线性插值替代传统线性混合,捕捉隐空间非线性结构。
- 在不依赖微调条件下,实现多样风格的高保真一致融合。
- 适合需要快速多风格生成且避免训练成本的研究者。
我们提出 Z-SASLM,一种零样本风格对齐的 Spherical Linear Interpolation(SLI)混合隐空间操控框架,克服了现有多风格混合方法的局限。传统方法依赖线性混合,假设隐空间为平坦结构,导致多参考风格融合效果不佳。相比之下,本框架利用隐空间的非线性几何特性,通过 SLI 混合加权风格表示,沿超球面测地线进行插值,保留隐空间的内在结构,确保多样风格融合的高保真与连贯性,且无需微调。我们还提出新的评估指标 Weighted Multi-Style DINO ViT-B/8,用于定量衡量融合风格的一致性。尽管重点在于 SLI 混合在风格操控中的理论与实践优势,我们也通过全面实验验证其在多模态内容融合场景的有效性。结果表明,Z-SASLM 实现了增强且稳健的风格对齐。代码已公开于:https://github.com/alessioborgi/Z-SASLM。
原文摘要 · Abstract (English)
We introduce Z-SASLM, a Zero-Shot Style-Aligned SLI (Spherical Linear Interpolation) Blending Latent Manipulation pipeline that overcomes the limitations of current multi-style blending methods. Conventional approaches rely on linear blending, assuming a flat latent space leading to suboptimal results when integrating multiple reference styles. In contrast, our framework leverages the non-linear geometry of the latent space by using SLI Blending to combine weighted style representations. By interpolating along the geodesic on the hypersphere, Z-SASLM preserves the intrinsic structure of the latent space, ensuring high-fidelity and coherent blending of diverse styles - all without the need for fine-tuning. We further propose a new metric, Weighted Multi-Style DINO ViT-B/8, designed to quantitatively evaluate the consistency of the blended styles. While our primary focus is on the theoretical and practical advantages of SLI Blending for style manipulation, we also demonstrate its effectiveness in a multi-modal content fusion setting through comprehensive experimental studies. Experimental results show that Z-SASLM achieves enhanced and robust style alignment. The implementation code can be found at: https://github.com/alessioborgi/Z-SASLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。