arXiv:2601.17697cs.CV2026-01被引 3

用无风格参考特征分离艺术风格,无需微调即可实现通用风格解耦。

StyleDecoupler: Generalizable Artistic Style Disentanglement

  • 利用单模态模型提取纯内容特征,作为参考来剥离多模态嵌入中的风格信息。
  • 在WeART和WikiART上达到当前最佳风格检索性能,支持风格关系映射与生成评估。
  • 适合作为插件模块集成到冻结的视觉语言模型中,适合风格分析与生成研究者使用。

艺术风格的表示因与语义内容深度纠缠而极具挑战。我们提出StyleDecoupler,一种基于信息论的框架,核心思路是:多模态视觉模型同时编码风格与内容,而单模态模型则抑制风格以聚焦内容不变特征。通过将单模态表示作为仅含内容的参考,我们利用互信息最小化从多模态嵌入中分离出纯净风格特征。StyleDecoupler作为即插即用模块运行于冻结的视觉-语言模型之上,无需微调。我们还构建了WeART,一个包含280万幅作品、152种风格和1,556位艺术家的大规模基准数据集。实验表明,在WeART和WikiART上的风格检索任务中达到顶尖性能,同时支持风格关系映射与生成模型评估。方法与数据集已公开。

原文摘要 · Abstract (English)

Representing artistic style is challenging due to its deep entanglement with semantic content. We propose StyleDecoupler, an information-theoretic framework that leverages a key insight: multi-modal vision models encode both style and content, while uni-modal models suppress style to focus on content-invariant features. By using uni-modal representations as content-only references, we isolate pure style features from multi-modal embeddings through mutual information minimization. StyleDecoupler operates as a plug-and-play module on frozen Vision-Language Models without fine-tuning. We also introduce WeART, a large-scale benchmark of 280K artworks across 152 styles and 1,556 artists. Experiments show state-of-the-art performance on style retrieval across WeART and WikiART, while enabling applications like style relationship mapping and generative model evaluation. We release our method and dataset at this url.

风格解耦视觉语言模型艺术生成无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。