构建首个评估图像美学适宜性的大规模数据集与基准。
AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

- 提出双组件数据集:多维度美学批评与真实场景适用性评估
- 54,300张图像生成519,136条批评指令,301个专家评审场景
- 发现专精美学模型在情境判断上显著落后于通用模型
多模态大语言模型(MLLMs)已将图像美学评估从单一评分拓展至可解释的批评与指导。然而现有基准主要关注内在视觉质量或固定领域标准,未解决美观图像是否适合特定用途、受众、文化背景或领域规范的问题。我们提出AesCanvas,一个包含两个互补组件的统一评测体系:CritiqueCanvas涵盖54,300张图像的519,136条指令-响应对,支持摄影、绘画与虚拟图像的长篇多维度批评;ContextCanvas包含301个专家评审的真实使用场景,评估图像在具体情境中的美学适宜性。在统一协议下,我们评估了闭源前沿模型、开源通用模型及美学专用模型。结果揭示批评生成与情境敏感判断之间存在明显差异:基于参考的词法与语义指标仅部分捕捉批评质量,而美学专用模型在部分批评指标上仍具竞争力,但在ContextCanvas上显著落后于强通用模型。进一步分析表明,美学专长无法可靠迁移至情境适宜性,且模型决策可能未能追踪或扎根于关键情境视觉线索。这些发现确立了文化情境化、证据支撑的适宜性作为美学建模的独立目标。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving open whether an appealing image is appropriate for a specific purpose, audience, cultural setting, or domain convention. We introduce AesCanvas, a unified suite with two complementary components: CritiqueCanvas with 519,136 instruction-response pairs from 54,300 images supports long-form, multi-dimensional critique across photography, painting, and virtual imagery, whereas ContextCanvas with 301 expert-reviewed use scenarios evaluates contextual aesthetic suitability in realistic use scenarios. Under a unified protocol, we evaluate closed-source frontier, open-weight general, and aesthetic-specific MLLMs. Results reveal a clear separation between critique generation and context-sensitive judgment: reference-based lexical and semantic metrics only partially capture critique quality, while aesthetic specialists remain competitive on selected critique metrics yet substantially lag strong general-purpose MLLMs on ContextCanvas. Further analyses show that aesthetic specialization does not reliably transfer to contextual suitability and that model decisions may fail to track or ground themselves in decisive contextual visual cues. These findings establish culturally situated, evidence-grounded suitability as a distinct objective for aesthetic modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。