arXiv:2604.19587cs.CV2026-04

无需指令,自动识别并修复照片缺陷,提升视觉美感。

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

论文配图:SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
图 1 · 摘自论文原文
  • 将修图拆解为推理-生成一体化流程,自动理解图像问题
  • 在多个数据集上优于现有模型,对光影调整更敏感精准
  • 适合普通用户快速美化照片,也适用于专业级图像优化

传统摄影修图依赖用户具备美学知识来提供具体调整指令,但此类指令常模糊、不完整或非专业人士难以表达。本文提出 SmartPhotoCrafter,一种自动摄影图像编辑方法,将修图建模为紧密耦合的推理-生成过程。首先通过图像批评模块理解图像质量并识别缺陷,再由摄影艺术家模块针对性地增强画面吸引力,无需显式人类指令。采用多阶段训练:(i) 基础预训练建立基本审美与编辑能力;(ii) 以推理引导的多编辑监督进行适配,融入丰富语义指导;(iii) 推理-生成协同强化学习,联合优化推理与生成。训练中强调真实感生成,同时支持图像修复与修饰任务,并保持色彩与色调语义一致性。构建了阶段性数据集,逐步实现推理与可控生成、跨模块协作,最终达成高质量摄影增强。实验表明,SmartPhotoCrafter 在自动摄影增强任务中优于现有生成模型,生成结果更逼真,且对调色指令表现出更高色调敏感度。

原文摘要 · Abstract (English)

Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and camera parameters. However, this paradigm relies on explicit human instruction of aesthetic intent, which is often ambiguous, incomplete, or inaccessible to non-expert users. In this work, we propose SmartPhotoCrafter, an automatic photographic image editing method which formulates image editing as a tightly coupled reasoning-to-generation process. The proposed model first performs image quality comprehension and identifies deficiencies by the Image Critic module, and then the Photographic Artist module realizes targeted edits to enhance image appeal, eliminating the need for explicit human instructions. A multi-stage training pipeline is adopted: (i) Foundation pretraining to establish basic aesthetic understanding and editing capabilities, (ii) Adaptation with reasoning-guided multi-edit supervision to incorporate rich semantic guidance, and (iii) Coordinated reasoning-to generation reinforcement learning to jointly optimize reasoning and generation. During training, SmartPhotoCrafter emphasizes photo-realistic image generation, while supporting both image restoration and retouching tasks with consistent adherence to color- and tone-related semantics. We also construct a stage-specific dataset, which progressively builds reasoning and controllable generation, effective cross-module collaboration, and ultimately high-quality photographic enhancement. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models on the task of automatic photographic enhancement, achieving photo-realistic results while exhibiting higher tonal sensitivity to retouching instructions. Project page: https://github.com/vivoCameraResearch/SmartPhotoCrafter.

图像修复自动修图美学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。