用生成模型创造真实图像修复的高质量训练数据,解决数据稀缺难题。
GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

- 用多模态大模型生成真实低质图对应的高质量目标图
- 构建含10.37万对图像的GGT-100K数据集,覆盖复杂真实退化
- 显著提升各类修复模型在真实场景下的泛化能力,尤其适合生成式模型微调
真实世界图像修复(IR)受限于高质量成对训练数据的匮乏。合成数据虽丰富但难以模拟真实退化,而真实成对数据成本高且难获取。为此,本文提出生成式真值(GGT),利用多模态基础模型(MFMs)从真实低质量(LQ)图像生成高质量(HQ)目标图。我们系统评估了九个先进MFMs(包括Nano-Banana-2和GPT-Image-2),发现基于视觉语言模型的自适应提示策略使Nano-Banana-2在生成感知真实且内容忠实的高清图像方面表现最优。基于此,构建了包含103,707个训练对、500个测试对的GGT-100K数据集,涵盖多样场景与复杂真实退化。大量实验表明,该数据集能持续提升多种IR模型的真实世界泛化性能,尤其利于生成式模型的微调。结果表明,MFMs可作为修复导向数据生成的有效工具,GGT-100K是拓展真实世界修复模型泛化边界的宝贵资源。
原文摘要 · Abstract (English)
Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail to model real-world degradations, while real-world paired datasets are expensive and difficult to capture. As a result, IR models trained on these datasets show limited generalization in real-world scenarios. In this work, we propose Generative Ground Truth (GGT) by using generative multimodal foundation models (MFMs) to produce high-quality (HQ) targets from real-world low-quality (LQ) images. We first conduct a systematic evaluation of nine state-of-the-art MFMs, including Nano-Banana-2 and GPT-Image-2, on images of various scenes and degradation types. The results demonstrate that Nano-Banana-2 with VLM-based adaptive prompting shows the highest capability to synthesize perceptually realistic and content-faithful HQ targets, which can serve as the GGT for the LQ input. We then employ Nano-Banana-2 to build a GGT synthesis pipeline, which involves multi-stage quality control to ensure data reliability, and construct GGT-100K, an LQ-HQ paired dataset comprising 103,707 training pairs and covering diverse scenes and complex real-world degradations. A test set of 500 image pairs is also established. Extensive experiments show that GGT-100K consistently improves the real-world generalization of a wide range of IR models, with particularly strong benefits for finetuning generative models for IR tasks. Our results suggest that MFMs can serve as practical tools for restoration-oriented data generation, and GGT-100K is a useful resource to expand the generalization boundaries of real-world IR models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。