让图像超分模型更懂人眼偏好,实时优化生成质量。
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
- 用动态权衡机制在线评估低清到高清的转换质量。
- 在真实图像超分任务中同时提升视觉美感与保真度。
- 适合需要高质量图像生成的实时应用或内容创作。
由于感知与保真度之间的权衡以及未知的复杂退化,使生成式真实世界图像超分辨率模型与人类视觉偏好对齐极具挑战。现有方法依赖离线偏好优化和静态指标聚合,往往不可解释且在强条件约束下易产生伪多样性。本文提出OARS,一种基于COMPASS(一种基于多模态大模型的奖励模型)的流程感知在线对齐框架,该模型通过联合建模保真度保持与感知增益,并采用输入质量自适应的权衡策略评估从低质到高质图像的转换过程。为训练COMPASS,我们构建了涵盖合成与真实退化的COMPASS-20K数据集,并设计三阶段感知标注流程,获得校准且细粒度的训练标签。OARS在浅层LoRA优化支持下,从冷启动流匹配逐步推进至全参考与无参考强化学习,实现在线策略探索。大量实验与用户研究证明,该方法在保持保真度的同时持续提升感知质量,在Real-ISR基准上达到当前最优性能。
原文摘要 · Abstract (English)
Aligning generative real-world image super-resolution models with human visual preference is challenging due to the perception--fidelity trade-off and diverse, unknown degradations. Prior approaches rely on offline preference optimization and static metric aggregation, which are often non-interpretable and prone to pseudo-diversity under strong conditioning. We propose OARS, a process-aware online alignment framework built on COMPASS, a MLLM-based reward that evaluates the LR to SR transition by jointly modeling fidelity preservation and perceptual gain with an input-quality-adaptive trade-off. To train COMPASS, we curate COMPASS-20K spanning synthetic and real degradations, and introduce a three-stage perceptual annotation pipeline that yields calibrated, fine-grained training labels. Guided by COMPASS, OARS performs progressive online alignment from cold-start flow matching to full-reference and finally reference-free RL via shallow LoRA optimization for on-policy exploration. Extensive experiments and user studies demonstrate consistent perceptual improvements while maintaining fidelity, achieving state-of-the-art performance on Real-ISR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。