用多尺度文本引导自监督学习,提升图像审美评估的全面性与个性化。
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
- 设计多尺度特征对齐机制,利用无标签数据自监督训练。
- 在审美评分、评论生成等任务上达到新最好性能。
- 支持零样本审美建议与个性化评估,适合内容创作场景。
图像美学评估(IAA)是一项关键而复杂的任务,需分析图像的美学价值并识别其亮点与改进点。传统方法通常聚焦单一任务且依赖有限标注数据,难以实现深度美学理解。尽管多模态大模型(MLLM)被尝试用于解决此问题,但在IAA领域仍不成熟。为此,我们提出一种具备细致审美洞察力的综合性MLLM,核心是创新的多尺度文本引导自监督学习技术。该技术包含多尺度特征对齐模块,通过大量无标签数据进行自监督训练,从结构与功能上增强模型的美学能力。实证表明,在广泛指令微调后,模型在多个任务(如美学打分、美学评论生成、个性化图像审美评估)中均达到新最优水平。尤为突出的是,它在新兴的美学建议任务中展现出零样本学习能力。此外,针对个性化评估,我们利用上下文学习潜力,验证了其内在优势。
原文摘要 · Abstract (English)
Image Aesthetic Assessment (IAA) is a vital and intricate task that entails analyzing and assessing an image's aesthetic values, and identifying its highlights and areas for improvement. Traditional methods of IAA often concentrate on a single aesthetic task and suffer from inadequate labeled datasets, thus impairing in-depth aesthetic comprehension. Despite efforts to overcome this challenge through the application of Multi-modal Large Language Models (MLLMs), such models remain underdeveloped for IAA purposes. To address this, we propose a comprehensive aesthetic MLLM capable of nuanced aesthetic insight. Central to our approach is an innovative multi-scale text-guided self-supervised learning technique. This technique features a multi-scale feature alignment module and capitalizes on a wealth of unlabeled data in a self-supervised manner to structurally and functionally enhance aesthetic ability. The empirical evidence indicates that accompanied with extensive instruct-tuning, our model sets new state-of-the-art benchmarks across multiple tasks, including aesthetic scoring, aesthetic commenting, and personalized image aesthetic assessment. Remarkably, it also demonstrates zero-shot learning capabilities in the emerging task of aesthetic suggesting. Furthermore, for personalized image aesthetic assessment, we harness the potential of in-context learning and showcase its inherent advantages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。