让AI真正理解艺术家风格,避免用套路生成画作。
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

- 构建显式控制状态,分离场景、风格与避坑约束
- 在梵高和齐白石作品上显著提升风格保真度与结构一致性
- 适合需要精准艺术风格还原的研究者与创作者
艺术家风格化图像生成不能仅靠在提示中添加艺术家名字。当前模型常依赖固定套路,如重复元素、通用配色或典型时期特征,而非忠实还原用户意图。本文提出Atelier框架,将模糊的艺术意图转化为显式的控制状态,包含场景锚点、保留/转换决策、风格假设、角色绑定的艺术家证据及避坑约束。该状态通过艺术家知识与局部图像片段进行锚定,生成适配后端的生成计划,并通过全局与局部真实感反馈迭代优化。我们还构建了ArtIntentBench基准,涵盖梵高与齐白石在重绘、风格控制、历史未见主题、套路审计与人类偏好评估上的测试。在开源与闭源生成器上,Atelier均显著优于提示工程、检索增强与通用代理基线,在风格保真度与结构保留上更优,且大幅减少套路替代现象。结果表明,艺术风格生成的瓶颈不仅在于图像合成,更在于上游对显式、证据驱动的艺术控制的推断能力。
原文摘要 · Abstract (English)
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。