arXiv:2507.14533cs.CV2025-07被引 38

提出可同时打分与解析美学细节的AI图像评估模型。

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

  • 用多模态大模型实现评分与专家级美学分析联合输出。
  • 构建1万张专业标注的图像数据集,含8维属性分解。
  • 适合需要深度美学理解的AI艺术生成与教育应用。

教育应用、艺术创作及AI生成内容(AIGC)技术的快速发展,显著提升了对全面图像美学评估(IAA)的需求,尤其要求方法既能提供量化评分,又能具备专业理解能力。基于多模态大语言模型(MLLM)的IAA方法相比传统方法展现出更强的感知与泛化能力,但仍存在仅输出分数或仅输出文本的模态偏差,且缺乏细粒度属性分解,难以支持进一步的美学分析。本文提出:(1) ArtiMuse,一种具备联合评分与专家级理解能力的MLLM基图像美学评估模型;(2) ArtiMuse-10K,首个由专业专家标注的图像美学数据集,包含10,000张图像,覆盖5个主类别与15个子类别,每张图像均附有8维度属性分析与整体评分。模型与数据集将公开,以推动该领域发展。

原文摘要 · Abstract (English)

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly demanding methods capable of delivering both quantitative scoring and professional understanding. Multimodal Large Language Model (MLLM)-based IAA methods demonstrate stronger perceptual and generalization capabilities compared to traditional approaches, yet they suffer from modality bias (score-only or text-only) and lack fine-grained attribute decomposition, thereby failing to support further aesthetic assessment. In this paper, we present:(1) ArtiMuse, an innovative MLLM-based IAA model with Joint Scoring and Expert-Level Understanding capabilities; (2) ArtiMuse-10K, the first expert-curated image aesthetic dataset comprising 10,000 images spanning 5 main categories and 15 subcategories, each annotated by professional experts with 8-dimensional attributes analysis and a holistic score. Both the model and dataset will be made public to advance the field.

图像评估多模态美学分析AIGC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。