arXiv:2412.07152cs.CV2024-12被引 6

用一步扩散模型提升超分辨率,更符合人眼感知。

Hero-SR: One-Step Diffusion for Super-Resolution with Human Perception Priors

  • 引入动态时间步模块,自适应选择最佳生成步骤。
  • 结合图文双模态监督,提升语义一致性和自然感。
  • 适合需要真实感超分的图像修复与生成任务。

由于扩散模型的强大先验,近期方法在真实世界超分辨率(Real-SR)中展现出潜力。然而,在重度退化和复杂输入条件下,实现语义一致性与人眼感知自然性仍具挑战。为此,我们提出Hero-SR,一种基于一步扩散的超分辨率框架,显式融入人类感知先验。该框架包含两个新模块:动态时间步模块(DTSM),可自适应选择最优扩散步数以灵活匹配人眼感知标准;开放世界多模态监督(OWMS),通过CLIP融合图像与文本域指导,提升语义一致性和感知自然性。得益于这些设计,Hero-SR生成的高分辨率图像不仅保留精细细节,还符合人类感知偏好。大量实验表明,Hero-SR在Real-SR任务上达到当前最优性能。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Owing to the robust priors of diffusion models, recent approaches have shown promise in addressing real-world super-resolution (Real-SR). However, achieving semantic consistency and perceptual naturalness to meet human perception demands remains difficult, especially under conditions of heavy degradation and varied input complexities. To tackle this, we propose Hero-SR, a one-step diffusion-based SR framework explicitly designed with human perception priors. Hero-SR consists of two novel modules: the Dynamic Time-Step Module (DTSM), which adaptively selects optimal diffusion steps for flexibly meeting human perceptual standards, and the Open-World Multi-modality Supervision (OWMS), which integrates guidance from both image and text domains through CLIP to improve semantic consistency and perceptual naturalness. Through these modules, Hero-SR generates high-resolution images that not only preserve intricate details but also reflect human perceptual preferences. Extensive experiments validate that Hero-SR achieves state-of-the-art performance in Real-SR. The code will be publicly available upon paper acceptance.

超分辨率扩散模型人眼感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。