arXiv:2511.06360cs.CV2025-11被引 1

构建首个覆盖感知与创作的多模态审美评测基准

AesTest: Measuring Aesthetic Intelligence from Perception to Production

  • 设计十类任务,融合心理学理论评估审美能力
  • 整合专业流程与大众偏好数据,覆盖专家与真实场景
  • 支持属性分析、情感共鸣等多元审美查询,适合研究者使用

感知与生成审美判断是多模态大语言模型(MLLMs)的一项基础但未充分探索的能力。现有图像审美评估(IAA)基准在感知范围上狭窄或缺乏系统性审美生成评估所需多样性。为此,我们提出AesTest,一个全面的多模态审美感知与生成评测基准,具备以下特点:1)包含十个任务的精选选择题,涵盖感知、欣赏、创作与摄影,基于生成学习的心理学理论;2)整合来自专业编辑流程、摄影构图教程和众包偏好的数据,确保覆盖专家原则与真实世界多样性;3)支持多种审美查询类型,如属性分析、情感共鸣、构图选择与风格推理。我们在AesTest上评估了指令微调的IAA MLLMs与通用MLLMs,揭示了构建审美智能的重大挑战。我们将公开发布AesTest,以支持该领域的未来研究。

原文摘要 · Abstract (English)

Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aesthetic assessment (IAA) are narrow in perception scope or lack the diversity needed to evaluate systematic aesthetic production. To address this gap, we introduce AesTest, a comprehensive benchmark for multimodal aesthetic perception and production, distinguished by the following features: 1) It consists of curated multiple-choice questions spanning ten tasks, covering perception, appreciation, creation, and photography. These tasks are grounded in psychological theories of generative learning. 2) It integrates data from diverse sources, including professional editing workflows, photographic composition tutorials, and crowdsourced preferences. It ensures coverage of both expert-level principles and real-world variation. 3) It supports various aesthetic query types, such as attribute-based analysis, emotional resonance, compositional choice, and stylistic reasoning. We evaluate both instruction-tuned IAA MLLMs and general MLLMs on AesTest, revealing significant challenges in building aesthetic intelligence. We will publicly release AesTest to support future research in this area.

多模态审美评估大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。