让大模型更懂图片美感,精准生成审美描述。
Aesthetic Image Captioning with Saliency Enhanced MLLMs
- 引入视觉美感注意力模块,让模型聚焦关键美学区域。
- 在主流评测集上超越传统方法与通用大模型,达最新水平。
- 适合研究图像审美生成、多模态大模型优化的学者。
美学图像描述(Aesthetic Image Captioning, AIC)旨在生成描述图像美学特征的文本,是计算美学领域的重要方向。近年来,预训练多模态大语言模型(MLLMs)快速发展,推动了视觉与文本融合的美学研究。然而,现有研究多集中于美学评分预测,对AIC的应用有限。现有基于MLLM的AIC方法主要依赖微调,未专门适配模型关注目标美学内容。为此,本文提出美学显著性增强多模态大语言模型(ASE-MLLM),一个端到端框架,显式将美学显著性融入MLLM。该框架包含图像美学显著性模块(IASM),高效提取图像美学特征;并设计IAS-ViT作为图像编码器,通过交叉注意力机制融合美学特征与原始图像特征。据我们所知,ASE-MLLM是首个为AIC任务专门集成图像美学显著性的框架。大量实验表明,该方法在主流AIC基准上显著优于传统方法与通用MLLM,达到当前最优性能。
原文摘要 · Abstract (English)
Aesthetic Image Captioning (AIC) aims to generate textual descriptions of image aesthetics, becoming a key research direction in the field of computational aesthetics. In recent years, pretrained Multimodal Large Language Models (MLLMs) have advanced rapidly, leading to a significant increase in image aesthetics research that integrates both visual and textual modalities. However, most existing studies on image aesthetics primarily focus on predicting aesthetic ratings and have shown limited application in AIC. Existing AIC works leveraging MLLMs predominantly rely on fine-tuning methods without specifically adapting MLLMs to focus on target aesthetic content. To address this limitation, we propose the Aesthetic Saliency Enhanced Multimodal Large Language Model (ASE-MLLM), an end-to-end framework that explicitly incorporates aesthetic saliency into MLLMs. Within this framework, we introduce the Image Aesthetic Saliency Module (IASM), which efficiently and effectively extracts aesthetic saliency features from images. Additionally, we design IAS-ViT as the image encoder for MLLMs, this module fuses aesthetic saliency features with original image features via a cross-attention mechanism. To the best of our knowledge, ASE-MLLM is the first framework to integrate image aesthetic saliency into MLLMs specifically for AIC tasks. Extensive experiments demonstrated that our approach significantly outperformed traditional methods and generic MLLMs on current mainstream AIC benchmarks, achieving state-of-the-art (SOTA) performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。