arXiv:2608.18622cs.CV2026-08

让AI懂每个人的审美,用极小成本实现个性化修图。

PALATE: Personalized Aesthetic Learning through Adaptive Taste Evolution for Multi-User Portrait Retouching

论文配图:PALATE: Personalized Aesthetic Learning through Adaptive Taste Evolution for Multi-User Portrait Retouching
图 1 · 摘自论文原文
  • 用三层共享+轻量适配器,动态学习用户偏好。
  • 仅需少量评分就能预测用户喜好,准确率达72.83%。
  • 适合需要快速适配多用户的个性化修图场景。

自动人像精修发展迅速,但其目标本质上是主观的:同一张人像存在多种专业有效的处理结果,用户间对最佳效果意见不一。现有方法多基于全局美学标准优化,无法捕捉个体审美;而为每位用户单独微调模型则带来高昂的训练、存储与数据成本。本文提出PALATE,一种共享奖励演化框架,保持图像编辑器不变,转而个性化选择同一源图的不同修饰候选。该框架将用户奖励分解为全局共享主干、类别级残差(服务于审美相似用户)和轻量级用户适配器,并引入防坍塌正则化确保三者互补。采用循环双层蒸馏机制,先将用户特定偏好蒸馏至类别奖励,再将类别知识回传至全局主干,用于初始化下一轮演化。如此迭代,共享初始化逐步提升,使新用户仅凭少量排序即可完成校准。在包含10,000张专家精修候选图的PPR10K数据集上,对保留用户和保留图像进行测试,PALATE达到72.83%的成对偏好预测准确率,显著优于所有奖励、美学及图像质量基线,其中最强基线PickScore为58.06%。每位新用户仅需512字节参数,评分耗时毫秒级。

原文摘要 · Abstract (English)

Automatic portrait retouching has advanced rapidly, yet its objective is inherently subjective: the same portrait admits multiple professionally valid results, and users disagree about which one is best. Most existing methods optimize a population-level aesthetic standard and therefore cannot capture individual taste, while fine-tuning a separate editing model for every user incurs prohibitive training, storage, and data costs. We propose PALATE, a shared reward-evolution framework that keeps the image editor fixed and instead personalizes the selection among retouched candidates of the same source portrait. PALATE decomposes the reward for each user into a global backbone shared by all users, category-level residuals shared by aesthetically similar users, and a lightweight user adapter, with anti-collapse regularizers keeping the three levels complementary.A cyclic dual-level distillation scheme first distills user-specific preferences into category rewards and then consolidates the resulting category-level knowledge into the global backbone, which is redistributed to initialize the next evolution round. In this way, the shared initialization improves progressively across rounds, enabling unseen users to be calibrated from only a few rankings. On expert-retouched candidates from PPR10K with held-out users and held-out images, PALATE attains 72.83% pairwise preference-prediction accuracy, surpassing all reward, aesthetic, and image-quality baselines, of which the strongest, PickScore, reaches 58.06%. Each new user costs only 512 bytes of user-specific parameters and millisecond-level scoring.

个性化修图美学学习轻量化适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。