arXiv:2605.12545cs.CVcs.AI2026-05

让AI像专业摄影师一样思考,自动裁剪出更符合审美标准的图片。

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference

论文配图:CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference
图 1 · 摘自论文原文
  • 将图像裁剪转化为分步推理任务,模拟摄影师的创作逻辑。
  • 在多个数据集上显著优于现有方法,更贴近人类专家的裁剪偏好。
  • 适合需要高质量视觉输出的设计、摄影与内容生成场景。

美学图像裁剪旨在通过空间裁剪提升图像的构图美感。以往方法多依赖显著性预测或检索增强,忽视了该任务的核心:对构图与审美的深层理解。因此,基于显著性的方法难以在复杂场景中做出合理的构图权衡,而基于检索的方法盲目参考相似案例,缺乏对独特场景的自适应推理能力,导致自动化裁剪结果与人类专家不一致。为解决上述问题,我们提出一种新范式,将美学裁剪重新定义为多模态推理任务,旨在激活视觉语言模型(VLM)在美学方面的分析与理解能力。我们设计了基于组合推理与偏好优化的CROP方法,引导VLM像专业摄影师一样思考:通过分解场景元素与构图原则,逐步完成‘分析-提议-决策’流程。同时,专家偏好对齐模块使模型决策更符合人类专家的审美标准。在多个数据集上的大量实验验证了该方法的优越性及各组件的有效性。

原文摘要 · Abstract (English)

Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the task's core requirement: a deep understanding of composition and aesthetics. Consequently, saliency-based methods struggle to make compositional trade-offs in complex scenes, while retrieval-based methods blindly refer to similar cases, lacking adaptive reasoning for unique scenes. Both approaches fail to align their automated cropping results with those of human experts. To address the above issues, we propose a novel paradigm that reformulates aesthetic cropping as a multimodal reasoning task, aiming to activate the VLM's analytical and comprehension capabilities in aesthetics. We design a Compositional Reasoning and Optimizing Preference method (CROP) that directs the VLM to think like a professional photographer. It deconstructs a complex and subjective aesthetic problem into an "analysis-proposal-decision" process, reasoning step by step through the analysis of scene elements and compositional principles. Meanwhile, our expert preference alignment module makes the model's decision consistent with human expert aesthetics. Extensive experiments across multiple datasets validate our method's superiority and component effectiveness.

图像裁剪视觉推理审美生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。