arXiv:2602.22809cs.CV2026-02

用大模型自动规划美学编辑步骤,无需用户逐条指令

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

  • 将图像编辑视为长程决策问题,通过树搜索规划多步操作
  • 在7000张UGC图片上验证,比基线方法更准确遵循指令且视觉质量更高
  • 适合需要自动化修图的设计师或普通用户,尤其擅长复杂美学调整

随着生成模型快速发展,基于指令的图像编辑在生成高质量图像方面展现出巨大潜力。然而,编辑质量高度依赖精心设计的指令,导致任务分解与排序完全由用户承担。为实现自主图像编辑,我们提出PhotoAgent系统,通过显式美学规划推进图像编辑。具体而言,PhotoAgent将自主图像编辑建模为长程决策问题,推理用户审美意图,通过树搜索规划多步编辑动作,并借助记忆与视觉反馈进行闭环执行与迭代优化,无需用户逐步提示。为支持真实场景下的可靠评估,我们构建了UGC-Edit基准,包含7000张照片和一个学习得到的美学奖励模型,并建立包含1017张照片的测试集以系统评估自主图像编辑性能。大量实验表明,与基线方法相比,PhotoAgent在指令遵循度和视觉质量上均持续提升。

原文摘要 · Abstract (English)

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on the user. To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. Specifically, PhotoAgent formulates autonomous image editing as a long-horizon decision-making problem. It reasons over user aesthetic intent, plans multi-step editing actions via tree search, and iteratively refines results through closed-loop execution with memory and visual feedback, without requiring step-by-step user prompts. To support reliable evaluation in real-world scenarios, we introduce UGC-Edit, an aesthetic evaluation benchmark consisting of 7,000 photos and a learned aesthetic reward model. We also construct a test set containing 1,017 photos to systematically assess autonomous photo editing performance. Extensive experiments demonstrate that PhotoAgent consistently improves both instruction adherence and visual quality compared with baseline methods. The project page is https://mdyao.github.io/PhotoAgent/.

图像编辑美学规划大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。