AesCrop通过构图引导实现更美观的图像裁剪,兼顾全局性与多样性。
AesCrop: Aesthetic-driven Cropping Guided by Composition
- 融合评估与回归的混合框架,用新注意力机制注入构图先验。
- 在多个数据集上优于当前最优方法,质量评分提升3.2%以上。
- 适合需要高质量缩略图的应用,如视频推荐与内容展示。
基于美学的图像裁剪在视图推荐和缩略图生成等应用中至关重要,视觉吸引力显著影响用户参与度。视觉吸引力的关键因素是构图——图像元素的精心布局。现有方法通过评估式和回归式范式引入构图知识,但前者缺乏全局性,后者缺乏多样性。近期出现的混合方法虽弥补了这一差距,但在摄影构图引导方面仍存在不足。本文提出AesCrop,一种结合构图感知的混合图像裁剪模型,采用VMamba图像编码器并引入新型Mamba构图注意力偏置(MCAB),配合Transformer解码器,实现端到端基于排名的图像裁剪,可生成多个裁剪区域及对应质量评分。通过显式将构图线索编码至注意力机制,MCAB引导模型关注最具有构图意义的区域。大量实验表明,AesCrop在定量指标上超越当前最先进方法,定性结果也更符合审美偏好。
原文摘要 · Abstract (English)
Aesthetic-driven image cropping is crucial for applications like view recommendation and thumbnail generation, where visual appeal significantly impacts user engagement. A key factor in visual appeal is composition--the deliberate arrangement of elements within an image. Some methods have successfully incorporated compositional knowledge through evaluation-based and regression-based paradigms. However, evaluation-based methods lack globality while regression-based methods lack diversity. Recently, hybrid approaches that integrate both paradigms have emerged, bridging the gap between these two to achieve better diversity and globality. Notably, existing hybrid methods do not incorporate photographic composition guidance, a key attribute that defines photographic aesthetics. In this work, we introduce AesCrop, a composition-aware hybrid image-cropping model that integrates a VMamba image encoder, augmented with a novel Mamba Composition Attention Bias (MCAB) and a transformer decoder to perform end-to-end rank-based image cropping, generating multiple crops along with the corresponding quality scores. By explicitly encoding compositional cues into the attention mechanism, MCAB directs AesCrop to focus on the most compositionally salient regions. Extensive experiments demonstrate that AesCrop outperforms current state-of-the-art methods, delivering superior quantitative metrics and qualitatively more pleasing crops.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。