用大模型生成艺术图像审美数据,提升评估精度与效率
Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
- 构建70千条多维度审美描述数据集,降低人工标注成本
- 通过大模型解码实现长文本语义建模,性能超越现有方法
- 适合关注AIGC美学评估、跨模态学习的研究者使用
艺术图像审美质量评估对构建以人为本的AIGC量化评价体系至关重要。然而其涉及视觉感知、认知与情感等复杂层面,带来根本性挑战。现有数据集过度聚焦视觉感知,缺乏深层维度,且标注成本高;现有模型多采用多分支编码器分离审美属性,或依赖对比学习处理长文本描述,效果受限。为此,我们提出迭代生成的精炼审美描述(RAD)数据集,规模达70,000条,结构化支持多维度评估。同时设计ArtQuant框架,通过联合生成描述耦合审美维度,并利用大语言模型解码有效建模长文本语义。理论分析表明,RAD的数据语义充分性与生成范式协同降低预测熵,提供数学依据。该方法在多个数据集上达到最先进性能,仅需常规训练周期的33%,显著缩小艺术图像与审美判断之间的认知差距。代码与数据集将公开。
原文摘要 · Abstract (English)
The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature, spanning visual perception, cognition, and emotion, poses fundamental challenges. Although aesthetic descriptions offer a viable representation of this complexity, two critical challenges persist: (1) data scarcity and imbalance: existing dataset overly focuses on visual perception and neglects deeper dimensions due to the expensive manual annotation; and (2) model fragmentation: current visual networks isolate aesthetic attributes with multi-branch encoder, while multimodal methods represented by contrastive learning struggle to effectively process long-form textual descriptions. To resolve challenge (1), we first present the Refined Aesthetic Description (RAD) dataset, a large-scale (70k), multi-dimensional structured dataset, generated via an iterative pipeline without heavy annotation costs and easy to scale. To address challenge (2), we propose ArtQuant, an aesthetics assessment framework for artistic images which not only couples isolated aesthetic dimensions through joint description generation, but also better models long-text semantics with the help of LLM decoders. Besides, theoretical analysis confirms this symbiosis: RAD's semantic adequacy (data) and generation paradigm (model) collectively minimize prediction entropy, providing mathematical grounding for the framework. Our approach achieves state-of-the-art performance on several datasets while requiring only 33% of conventional training epochs, narrowing the cognitive gap between artistic images and aesthetic judgment. We will release both code and dataset to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。