让图像超分模型按需求选择真实或美观修复风格
FoA-SR: Faithful or Aesthetic? Profile-Aware Preference Optimization for Real-World Image Super-Resolution

- 基于用户偏好构建真实与美观两种修复风格的对比优化
- 在RealSR和DIV2K数据集上实现风格可控的高质量修复
- 通过独立奖励机制实现风格解耦,避免折中妥协
真实世界图像超分辨率常采用单一恢复目标,尽管生成模型可为同一输入生成多个高质量重建。本文提出FoA-SR,一种基于修复风格的偏好优化方法。首先训练基于FLUX.2的监督式超分适配器(Flux2SR),结合低分辨率隐空间条件、流匹配与图像空间重建损失。随后,为每张输入图生成共享随机候选池,并利用特定于风格的忠实与美学奖励对候选进行排序,挖掘胜者-败者对。这些对用于微调独立的LoRA适配器,保持基础模型冻结。在RealSR和DIV2K上的实验表明,该方法可将同一超分适配器导向不同恢复目标:忠实适配器提升参考一致性指标,美学适配器增强无参考感知质量指标。候选分析显示,两类奖励常选择不同胜者;混合LoRA消融实验表明,合并为单一奖励仅产生隐式折中,而非显式风格控制。
原文摘要 · Abstract (English)
Real-world image super-resolution (SR) is often designed with a single restoration objective, despite the current capacity of generative models to produce multiple high-quality reconstructions for the same input. In this paper, we argue that the best restoration strategy is subject to the specific restoration profile: a Faithful restoration prioritizes reference consistency, structure preservation, and hallucination suppression, whereas an Aesthetic restoration prioritizes visually pleasing and natural-looking details. We propose FoA-SR, a novel preference optimization approach to real-world SR based on profiles. To achieve this goal, FoA-SR starts with our supervised FLUX.2-based SR adapter (Flux2SR) trained with LR latent conditioning, flow matching, and image-space reconstruction losses for paired LR-to-HR image super-resolution. Following the development of the shared supervised super-resolution adapter, FoA-SR generates a shared stochastic candidate pool for each input image and ranks the same candidates using profile-specific Faithful and Aesthetic rewards to mine winner-loser pairs. These pairs are used to fine-tune separate LoRA adapters while keeping the base model frozen. Experiments on RealSR and DIV2K show that FoA-SR can steer the same SR adapter towards distinct restoration objectives: a Faithful adapter improves reference-consistent metrics while an Aesthetic adapter boosts metrics that measure perceptual quality without reference. Our candidate-pool analysis shows that Faithful and Aesthetic rewards frequently select different winners, and a Hybrid-LoRA ablation shows that collapsing both profiles into one reward yields an implicit compromise rather than explicit profile control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。