教大模型像摄影师一样看懂照片美感,提升专业级图像审美理解能力。
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- 构建多视角视觉融合机制,结合语言引导理解照片美学
- 在专业评测集上显著超越现有模型,实现更精准的美学分析
- 适合需要深度图像审美理解的摄影、设计与内容创作场景
摄影师在实际拍摄中常难以同时关注画面中的蓝色与天空,这揭示了通用视觉理解(如识别物体)与美学视觉理解(如将天空视为纯色块)之间的根本差异。这一区别给多模态大模型带来了挑战。尽管已有研究尝试探索,但大多局限于基础美学常识,难以应对真实场景中所需的摄影技术、前期/后期处理等专业知识。为此,我们提出首个基于专业摄影师与爱好者讨论的大规模、高多样性数据集PhotoCritique;设计新模型PhotoEye,采用语言引导的多视角视觉融合机制,从多角度理解图像美学;并构建全新基准测试PhotoBench,用于评估美学视觉理解能力。在多个现有基准及PhotoBench上,本模型均显著优于现有方法。
原文摘要 · Abstract (English)
While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky. Photographer and curator, Szarkowski insightfully revealed one of the notable gaps between general and aesthetic visual understanding: while the former focuses on identifying the factual element in an image (sky), the latter transcends such object identification, viewing it instead as an aesthetic component--a pure color block (blue). Such fundamental distinctions between general (detection, localization, etc.) and aesthetic (color, lighting, composition, etc.) visual understanding present a significant challenge for Multimodal Large Language Models (MLLMs). Although some recent works have made initial explorations, they are often limited to general and basic aesthetic commonsense. As a result, they frequently fall short in real-world scenarios (Fig. 1), which require extensive expertise--including photographic techniques, photo pre/post-processing knowledge, and more, to provide a detailed analysis and description. To fundamentally enhance the aesthetics understanding of MLLMs, we first introduce a novel dataset, PhotoCritique, derived from extensive discussions among professional photographers and enthusiasts, and characterized by the large scale, expertise, and diversity. Then, to better learn visual aesthetics from PhotoCritique, we furthur propose a novel model, PhotoEye, featuring a languageguided multi-view vision fusion mechanism to understand image aesthetics from multiple perspectives. Finally, we present a novel benchmark, PhotoBench, a comprehensive and professional benchmark for aesthetic visual understanding. On existing benchmarks and PhotoBench, our model demonstrates clear advantages over existing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。