评估并提升超分辨率模型的语义保真度,防止生成内容失真。
Evaluating and Preserving High-level Fidelity in Super-Resolution
- 构建首个带语义保真度标注的数据集,评估SR模型内容忠实性。
- SOTA模型在高阶保真度上表现不佳,视觉质量高但内容易错。
- 用基础模型优化可同时提升语义准确性和视觉质量,适合模型评估与改进。
近期图像超分辨率(SR)模型在细节重建和视觉效果上取得显著进展,但其强大的生成能力可能导致幻觉,改变图像内容,尽管视觉质量高。这种高层级内容变化人类容易识别,但现有低层级图像质量指标未能充分关注。本文强调测量高阶保真度的重要性,作为评估生成式SR模型可靠性的补充标准。我们构建了首个包含不同SR模型保真度评分的标注数据集,并评估了当前最先进(SOTA)SR模型在保持高阶保真度上的实际表现。基于该数据集,分析了现有图像质量指标与保真度测量的相关性,进一步表明此任务可通过基础模型更好解决。通过基于保真度反馈微调SR模型,我们验证了语义保真度和感知质量均可提升,展示了所提标准在模型评估与优化中的潜力。论文将公开数据集、代码与模型。
原文摘要 · Abstract (English)
Recent image Super-Resolution (SR) models are achieving impressive effects in reconstructing details and delivering visually pleasant outputs. However, the overpowering generative ability can sometimes hallucinate and thus change the image content despite gaining high visual quality. This type of high-level change can be easily identified by humans yet not well-studied in existing low-level image quality metrics. In this paper, we establish the importance of measuring high-level fidelity for SR models as a complementary criterion to reveal the reliability of generative SR models. We construct the first annotated dataset with fidelity scores from different SR models, and evaluate how state-of-the-art (SOTA) SR models actually perform in preserving high-level fidelity. Based on the dataset, we then analyze how existing image quality metrics correlate with fidelity measurement, and further show that this high-level task can be better addressed by foundation models. Finally, by fine-tuning SR models based on our fidelity feedback, we show that both semantic fidelity and perceptual quality can be improved, demonstrating the potential value of our proposed criteria, both in model evaluation and optimization. We will release the dataset, code, and models upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。