arXiv:2509.17012cs.CVcs.LG2025-09被引 1

构建5000张文档图像质量评估数据集,提出融合版式特征的高效评分模型

DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality Assessment

  • 设计多层级特征融合模块,结合低层视觉与高层版式信息
  • 在DIQA-5000和OCR相关数据集上超越现有通用图像质量模型
  • 支持多维度评分(整体质量/清晰度/色彩保真),适合文档处理系统优化

文档图像质量评估(DIQA)是光学字符识别(OCR)、文档修复及图像处理系统评价的重要环节。本文提出主观评估数据集DIQA-5000,包含5,000张图像,由500张真实场景文档经多种增强技术生成,每张图像由15名受试者在整体质量、清晰度和色彩保真三个维度打分。同时,我们提出一种专用无参考DIQA模型,利用文档版式特征在低分辨率下保持质量感知,降低计算开销。针对图像质量受低层与高层特征共同影响的特点,设计特征融合模块以提取并整合多层级特征。为实现多维评分,模型采用独立的质量头预测各维度得分分布,从而学习不同质量特性。实验表明,该方法在DIQA-5000和另一个侧重于OCR准确率的文档图像数据集上均优于当前最先进的通用图像质量评估模型。

原文摘要 · Abstract (English)

Document image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we introduce a subjective DIQA dataset DIQA-5000. The DIQA-5000 dataset comprises 5,000 document images, generated by applying multiple document enhancement techniques to 500 real-world images with diverse distortions. Each enhanced image was rated by 15 subjects across three rating dimensions: overall quality, sharpness, and color fidelity. Furthermore, we propose a specialized no-reference DIQA model that exploits document layout features to maintain quality perception at reduced resolutions to lower computational cost. Recognizing that image quality is influenced by both low-level and high-level visual features, we designed a feature fusion module to extract and integrate multi-level features from document images. To generate multi-dimensional scores, our model employs independent quality heads for each dimension to predict score distributions, allowing it to learn distinct aspects of document image quality. Experimental results demonstrate that our method outperforms current state-of-the-art general-purpose IQA models on both DIQA-5000 and an additional document image dataset focused on OCR accuracy.

文档图像质量评估特征融合无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。