模型越复杂,越难生成纯色图像,因存在审美偏见。
Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?

- 提出分级服从性框架,衡量模型从随机生成到像素级确定性的能力。
- 设计首个系统性基准Violin,发现闭源模型在确定性任务上更优。
- 揭示生成复杂图像强弱与生成纯色图像能力间的反向关联。
生成式AI在复杂内容创作中已达到人类水平,但我们发现一个'简单悖论':模型能渲染复杂场景,却难以完成低熵的简单任务,如生成均匀纯色图像。我们认为这是由不可控的涌现能力导致的系统性缺陷。随着模型规模扩大,对美学和复杂性的强烈先验会压制确定性简洁性,形成'审美偏见',阻碍模型从数据模拟过渡到真正的智力抽象。为此,我们正式提出AI服从性概念,构建分层级框架(1至5级),评估模型从概率逼近到像素级确定性的能力。引入首个系统性基准Violin,通过颜色纯度、图像掩码和几何形状生成三类确定性任务测试第4级服从性。评估多个先进模型后发现,闭源模型普遍优于开源模型,且本基准表现与自然图像生成基准相关。本工作为实现人类指令与模型输出更好对齐提供了基础框架与工具。
原文摘要 · Abstract (English)
Recent advances in generative AI have shown human-level performance in complex content creation. However, we identify a "Paradox of Simplicity": models that can render complex scenes often fail at trivial, low-entropy tasks, such as generating a uniform pure color image. We argue this is a systemic failure related to uncontrollable emergent abilities. As models scale, strong priors for aesthetics and complexity override deterministic simplicity, creating an "aesthetic bias" that hinders the model's transition from data simulation to true intellectual abstraction. To better investigate this problem, we formalize the concept of AI Obedience, a hierarchical framework that grades a model's ability to transition from probabilistic approximation to pixel-level determinism (Levels 1 to 5). We introduce Violin, the first systematic benchmark designed to evaluate Level 4 Obedience through three deterministic tasks: color purity, image masking, and geometric shape generation. Using Violin, we evaluate several state-of-the-art models and reveal that closed-source models generally outperform open-source ones in deterministic precision. Interestingly, performance on our benchmark correlates with the benchmark in natural image generation. Our work provides a foundational framework and tools for achieving better alignment between human instructions and model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。