arXiv:2608.07570cs.CVcs.AI2026-08

提出可解释的美学裁剪新框架,让裁剪理由有据可依。

COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

论文配图:COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping
图 1 · 摘自论文原文
  • 基于构图建立裁剪-解释联合学习框架
  • 构建3.3万组数据集,支持结构化输出
  • 适合研究可解释AI与图像编辑的开发者

可解释的美学图像裁剪不仅需定位视觉愉悦的裁剪区域,还需说明原因。现有方法多将解释视为事后生成的文本,忽视了构图这一关键美学因素。本文将可解释裁剪重构为结构化的裁剪-构图-解释问题,提出COMEX基准,通过图像扩展和输入输出反转流水线构建,包含33,161个四元组(扩展图像、裁剪框、构图类别、基于构图的解释),支持裁剪定位、构图理解与解释生成的联合学习。进一步提出两阶段SFT+GRPO框架:监督微调建立结构化输出规范与基础裁剪能力,GRPO提升裁剪质量、构图预测与解释忠实度。在COMEX及已有基准上评测15个大视觉语言模型与现有裁剪方法,验证框架有效性与可迁移性,各项指标表现优异。

原文摘要 · Abstract (English)

Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.

可解释性图像裁剪构图分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。