构建首个大规模低级视觉任务评估基准,揭示生成模型在图像修复中的短板。
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

- 设计涵盖16类退化任务的LL-Bench基准,含2469张真实退化图与28919张修复结果
- 发现生成模型在像素级控制上表现差,且存在特定失败模式,传统方法更稳定
- 提出基于多模态大模型的LL-Score,更贴近人类对修复质量与幻觉的判断
大规模生成模型在图像生成与编辑任务中表现卓越,但在需要像素级控制的低级视觉任务中性能仍不充分。为此,我们提出LL-Bench,一个全面评估生成模型在低级视觉任务能力的基准。该基准包含2,469张真实世界退化图像,覆盖16类低级退化任务,以及由10个先进生成模型和21个传统修复模型产生的28,919张修复图像,并附有152,020条专家级成对人类偏好标注和28,334个质量评分。基于此,我们系统诊断了生成模型在各类任务中的性能边界与独特失败模式,对比传统代表性修复方法。进一步研究发现现有质量评估指标与人类评分存在显著差异。为此,我们提出基于多模态大模型的LL-Score,能同时捕捉修复质量与幻觉存在性。大量实验表明,LL-Score不仅优于现有图像质量评估指标,还可作为低级视觉任务中生成模型训练的潜在奖励模型。
原文摘要 · Abstract (English)
Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks. However, their performance in low-level vision tasks, which require pixel-wise control, remains insufficiently studied. To address this gap, we introduce \textbf{LL-Bench}, a comprehensive \textbf{Benchmark} for evaluating the capabilities of large-scale generative models on \textbf{L}ow-\textbf{L}evel vision tasks. The benchmark comprises 2,469 real-world degraded images covering 16 low-level degradation tasks, and 28,919 restored images produced by 10 state-of-the-art large-scale generative models and 21 conventional restoration models, which are annotated with 152,020 expert-level pairwise human preferences and 28,334 quality scores. Built upon LL-Bench, we present a systematic diagnosis that reveals the performance boundaries and unique failure modes of large-scale generative models across diverse low-level vision tasks, compared with conventional representative restoration approaches. Moreover, we investigate the effectiveness of current quality evaluation metrics on LL-Bench, which exhibit significant discrepancy with human ratings. To better align restored-image quality assessment with human preferences, we further propose \textbf{LL-Score}, an MLLM-based evaluator that captures both restoration quality and hallucination existence. Extensive experiments demonstrate that LL-score not only outperforms existing image quality assessment metrics, but also serves as a promising reward model for training generative models on low-level vision tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。