arXiv:2603.01163cs.CV2026-03中稿 · CVPR被引 1

用强化学习提升人脸精修美感,避免噪点且更贴合人眼审美。

BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling

  • 引入动态路径引导,稳定采样过程防止噪声积累。
  • 构建10,000条细粒度偏好数据集,精准捕捉审美差异。
  • 适合追求高保真与主观美感一致的人脸处理场景。

人脸精修需在消除细微瑕疵的同时保留独特面部特征,以提升整体美感。现有方法面临根本性权衡:监督学习受限于像素级标签模仿,难以捕捉复杂的主观审美偏好;而在线强化学习虽能对齐偏好,其随机探索机制与高保真需求冲突,常因累积随机漂移引入明显噪声。为此,我们提出BeautyGRPO,一种面向美学对齐的强化学习框架。构建了包含五个关键精修维度的FRPref-10K细粒度偏好数据集,并训练专用奖励模型以评估细微感知差异。为调和探索与保真度,提出动态路径引导(DPG):通过动态计算基于锚点的常微分方程路径,并在每步采样时重规划引导轨迹,有效纠正随机漂移,同时保持可控探索。大量实验表明,BeautyGRPO优于专业人脸精修方法及通用图像编辑模型,在纹理质量、瑕疵去除准确性及与人类审美偏好一致性方面均表现更优。

原文摘要 · Abstract (English)

Face retouching requires removing subtle imperfections while preserving unique facial identity features, in order to enhance overall aesthetic appeal. However, existing methods suffer from a fundamental trade-off. Supervised learning on labeled data is constrained to pixel-level label mimicry, failing to capture complex subjective human aesthetic preferences. Conversely, while online reinforcement learning (RL) excels at preference alignment, its stochastic exploration paradigm conflicts with the high-fidelity demands of face retouching and often introduces noticeable noise artifacts due to accumulated stochastic drift. To address these limitations, we propose BeautyGRPO, a reinforcement learning framework that aligns face retouching with human aesthetic preferences. We construct FRPref-10K, a fine-grained preference dataset covering five key retouching dimensions, and train a specialized reward model capable of evaluating subtle perceptual differences. To reconcile exploration and fidelity, we introduce Dynamic Path Guidance (DPG). DPG stabilizes the stochastic sampling trajectory by dynamically computing an anchor-based ODE path and replanning a guided trajectory at each sampling timestep, effectively correcting stochastic drift while maintaining controlled exploration. Extensive experiments show that BeautyGRPO outperforms both specialized face retouching methods and general image editing models, achieving superior texture quality, more accurate blemish removal, and overall results that better align with human aesthetic preferences.

人脸精修强化学习美学对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。