arXiv:2505.22663cs.CV2025-05

无需训练,单图生成风格化抽象图像,兼顾识别与艺术变形。

Training Free Stylized Abstraction

  • 通过视觉语言模型在推理时提取身份特征,动态调整结构恢复策略。
  • 多轮生成不需微调,支持多种抽象风格(如LEGO、South Park)且泛化性强。
  • 提出新评估指标StyleBench,更符合人类对抽象图像的感知判断。

风格化抽象在保持语义忠实的同时,生成视觉夸张但可识别的图像表示。与强调结构保真的图像到图像转换不同,该任务需选择性保留身份线索并接受风格偏离,尤其对分布外个体更具挑战。本文提出一种无需训练的框架,仅用单张图像即可生成风格化抽象:利用视觉语言模型在推理阶段提取身份相关特征,并引入一种跨域校正流反演策略,基于风格依赖先验重建结构。通过风格感知的时间调度机制,动态调整结构修复过程,实现高保真还原,同时尊重主体与风格。支持无需微调的多轮抽象感知生成。为评估此任务,构建了基于GPT的人类对齐度量StyleBench,适用于像素级相似性失效的抽象风格场景。在多种抽象风格(如LEGO、针织娃娃、South Park)上的实验表明,方法在未见身份与风格上具备强泛化能力,整个系统完全开源。

原文摘要 · Abstract (English)

Stylized abstraction synthesizes visually exaggerated yet semantically faithful representations of subjects, balancing recognizability with perceptual distortion. Unlike image-to-image translation, which prioritizes structural fidelity, stylized abstraction demands selective retention of identity cues while embracing stylistic divergence, especially challenging for out-of-distribution individuals. We propose a training-free framework that generates stylized abstractions from a single image using inference-time scaling in vision-language models (VLLMs) to extract identity-relevant features, and a novel cross-domain rectified flow inversion strategy that reconstructs structure based on style-dependent priors. Our method adapts structural restoration dynamically through style-aware temporal scheduling, enabling high-fidelity reconstructions that honor both subject and style. It supports multi-round abstraction-aware generation without fine-tuning. To evaluate this task, we introduce StyleBench, a GPT-based human-aligned metric suited for abstract styles where pixel-level similarity fails. Experiments across diverse abstraction (e.g., LEGO, knitted dolls, South Park) show strong generalization to unseen identities and styles in a fully open-source setup.

风格化抽象视觉语言模型无训练生成图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。