让AI像专业摄影师一样指出照片问题并给出改进建议
Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping
- 用渐进式提问训练模型理解审美问题
- 在10748张图上实现SOTA的美学裁剪效果
- 适合想提升拍照质量的普通用户和设计师
智能手机普及使摄影无处不在,但普通用户与专业摄影师之间仍存在明显差距:后者能在拍摄时识别审美问题并提供可操作建议。我们定义这一能力为美学指导(AG),这是计算美学中尚未充分探索的重要领域。现有多模态大模型多提供过度积极反馈,无法识别问题或给出具体建议。缺乏AG能力导致其难以发现干扰区域或优化构图平衡,也限制了美学裁剪(aesthetic cropping)——即通过重框改善照片构图的能力。为此,我们提出AesGuide,首个大规模AG数据集与基准,包含10,748张带审美评分、分析和指导的图片。基于此,我们构建Venus框架,分两阶段实现:先通过逐步复杂的审美问题赋予多模态大模型AG能力,再利用思维链(CoT)推理激活其美学裁剪潜力。大量实验表明,Venus显著提升AG能力,在美学裁剪任务上达到当前最优(SOTA)表现,支持可解释且交互式的美学优化,贯穿拍摄前后全过程。代码已开源。
原文摘要 · Abstract (English)
The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We define this capability as aesthetic guidance (AG) -- an essential but largely underexplored domain in computational aesthetics. Existing multimodal large language models (MLLMs) primarily offer overly positive feedback, failing to identify issues or provide actionable guidance. Without AG capability, they cannot effectively identify distracting regions or optimize compositional balance, thus also struggling in aesthetic cropping, which aims to refine photo composition through reframing after capture. To address this, we introduce AesGuide, the first large-scale AG dataset and benchmark with 10,748 photos annotated with aesthetic scores, analyses, and guidance. Building upon it, we propose Venus, a two-stage framework that first empowers MLLMs with AG capability through progressively complex aesthetic questions and then activates their aesthetic cropping power via CoT-based rationales. Extensive experiments show that Venus substantially improves AG capability and achieves state-of-the-art (SOTA) performance in aesthetic cropping, enabling interpretable and interactive aesthetic refinement across both stages of photo creation. Code is available at https://github.com/PKU-ICST-MIPL/Venus_CVPR2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。