构建大规模图像编码评估数据集,支持精细质量分析。
JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding

- 从2787张图中筛选70张,结合语义聚类与人工精修。
- 覆盖12种编码器、17种配置,生成9618张失真图像。
- 揭示学习型编码器的评测分歧,适合压缩算法研究者。
近年来,传统与基于学习的图像编码技术发展迅速,对支持精细质量评估的基准数据集需求增加,尤其是针对基于学习的图像压缩方法。本文提出评估图像编码2026(AIC2026),一个用于高保真图像压缩的大规模数据集。该数据集包含70张源图像,是从2,787个候选图像中通过语义聚类、客观质量评估(IQA)方法间的指标分歧以及人工检查和精修筛选得出。数据集覆盖8种传统编码器和4种基于学习的编码器,共17种编码配置。每张源图像使用7种编码器进行编码,每个源-编码器组合在20个感知间距的失真级别上提供解码图像,失真水平对应约0.2–4.0刚可察觉差异(JND)单位,基于ColorVideoVDP(CVVDP)度量估算,总计生成9,618张失真图像。这种细粒度采样支持对率失真行为及客观评估指标在细微质量差异下的表现分析。我们使用24种传统和12种基于学习的IQA方法进行了广泛客观分析,结果表明当前IQA方法在细微质量差异上存在显著分歧,尤其在基于学习的编码器引入的失真方面更为明显。完整数据集可通过 https://doi.org/10.18419/DARUS-6156 公开获取。
原文摘要 · Abstract (English)
Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment of compressed image quality, particularly for learning-based image compression methods. This paper introduces Assessment of Image Coding 2026 (AIC2026), a large-scale dataset for high-fidelity image compression containing 70 source images selected from 2,787 candidates using semantic clustering, inter-metric disagreement among objective image quality assessment (IQA) methods, and manual inspection and refinement. The dataset covers a wide range of compression artifacts produced by eight conventional and four learning-based codecs across 17 coding configurations. Each source image is encoded using seven codecs. For each source-codec pair, decoded images are provided at 20 perceptually spaced distortion levels, corresponding approximately to 0.2-4.0 just-noticeable difference (JND) units using the ColorVideoVDP (CVVDP) metric for distortion estimation, yielding 9,618 distorted images. This fine-grained sampling enables analysis of rate-distortion behavior and objective metric evaluation for subtle quality differences across a wide range of compression artifacts. We report an extensive objective analysis using 24 conventional and 12 learning-based IQA methods. The results show substantial disagreement among current IQA methods for fine-grained quality differences, particularly for artifacts introduced by learning-based codecs. The complete dataset is publicly available at https://doi.org/10.18419/DARUS-6156.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。