arXiv:2505.22490cs.CV2025-05AAAI被引 6

用专业摄影指导自动裁图,效果媲美人工标注。

ProCrop: Learning Aesthetic Image Cropping from Professional Compositions

  • 从专业照片中提取构图规律,融合查询图像特征进行裁剪
  • 在24.2万张弱标注数据上训练,性能超越已有方法
  • 适合研究图像美学与自动构图的开发者和设计师

图像裁剪对提升照片视觉吸引力和叙事力至关重要,但现有基于规则或数据驱动的方法常缺乏多样性或依赖标注数据。我们提出ProCrop,一种基于检索的方法,利用专业摄影作品引导裁剪决策。通过融合专业照片与查询图像的特征,ProCrop学习专业构图模式,显著提升性能。此外,我们构建了一个包含24.2万张弱标注图像的大规模数据集,通过外补绘制专业图像并迭代优化多样化裁剪方案。该数据集基于美学原则生成高质量裁剪建议,是当前最大的公开图像裁剪数据集。大量实验表明,ProCrop在监督与弱监督设置下均显著优于现有方法。尤其在新数据集上训练后,其性能超越以往弱监督方法,甚至接近全监督水平。代码与数据集将公开,以推动图像美学与构图分析研究。

原文摘要 · Abstract (English)

Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide cropping decisions. By fusing features from professional photographs with those of the query image, ProCrop learns from professional compositions, significantly boosting performance. Additionally, we present a large-scale dataset of 242K weakly-annotated images, generated by out-painting professional images and iteratively refining diverse crop proposals. This composition-aware dataset generation offers diverse high-quality crop proposals guided by aesthetic principles and becomes the largest publicly available dataset for image cropping. Extensive experiments show that ProCrop significantly outperforms existing methods in both supervised and weakly-supervised settings. Notably, when trained on the new dataset, our ProCrop surpasses previous weakly-supervised methods and even matches fully supervised approaches. Both the code and dataset will be made publicly available to advance research in image aesthetics and composition analysis.

图像裁剪美学评估数据集构建弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。