让普通照片自动修复构图缺陷,生成更美的记忆画面。
AesFormer: Transform Everyday Photos into Beautiful Memories

- 分两阶段处理:先分析美学维度,再按指令重构图像结构
- 在9071对图像上测试,显著提升照片美感且保留人物与场景原貌
- 适合想提升日常摄影质量的用户和图像编辑研究者
日常摄影中,许多具有美感的瞬间常因构图、视角或姿态等结构性问题而影响视觉效果,现有修图与人像增强方法难以解决。本文将美学照片重建(APR)定义为在保持主体身份和场景语义的前提下,通过结构重构提升照片美感。尽管近期图像编辑模型使这一目标成为可能,但普遍缺乏美学理解,导致编辑结果语义合理却审美不足。为此,我们提出AesFormer,一种两阶段框架,将美学规划与图像编辑解耦。第一阶段,美学动作模型(AesThinker)沿七个渐进摄影维度分析输入并输出可执行编辑动作;我们进一步采用GRPO-A以鼓励在多样化动作方案中广泛探索,超越单一监督微调。第二阶段,动作条件编辑器(AesEditor)根据这些动作进行结构化编辑。为支持APR,我们构建了基于视频的语料挖掘流程(VCMP),并建立包含9071对严格对齐(差、好)图像的基准数据集AesRecon。实验表明,AesFormer显著提升APR性能,且优于Nano Banana Pro。代码已公开于https://github.com/PKU-ICST-MIPL/AesFormer_ICML2026。
原文摘要 · Abstract (English)
In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that existing retouching and portrait enhancement methods cannot fix. We formulate Aesthetic Photo Reconstruction (APR) as improving a photo's aesthetic quality via structural reconstruction while preserving subject identity and scene semantics. Although recent advances in image editing models make APR feasible, they often lack aesthetic understanding, yielding edits that are semantically plausible yet aesthetically weak. To address this, we propose AesFormer, a two-stage framework that decouples aesthetic planning from image editing. In Stage 1, an aesthetic action model (AesThinker) analyzes the input along seven progressive photographic dimensions and outputs executable editing actions; we further apply GRPO-A to encourage broad exploration over diverse action plans beyond SFT. In Stage 2, an action-conditioned editor (AesEditor) performs structural edits guided by these actions. To support APR, we build a video-based corpus-mining pipeline (VCMP) and construct AesRecon, a benchmark of 9,071 strictly aligned (poor, good) image pairs. Experiments show that AesFormer substantially improves APR performance and is competitive with Nano Banana Pro. Code is available at https://github.com/PKU-ICST-MIPL/AesFormer_ICML2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。