构建多领域文本图像检测基准,揭示现有模型在结构化文本图上的检测短板。
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

- 构建覆盖6类场景的12095张GPT-Image-2生成图像数据集
- 发现现有检测器在结构化文本图像上性能差异大且易受压缩影响
- 提示需发展关注文本与版式语义的新型检测方法,适合安全与内容可信研究者
文本丰富的图像常包含隐私、交易或决策相关的关键信息。随着多模态图像生成模型日益能合成逼真的文字内容与布局设计,检测这类AI生成图像已成为保障数字信任与内容真实性的关键挑战。现有基准大多聚焦于以物体为中心的图像,对文本语义与版式组织至关重要的场景覆盖不足。本文提出TextRich,一个针对OpenAI GPT-Image-2生成的文本丰富图像的多领域检测基准。该基准涵盖商业海报、信息图表、学术海报、收据、表格和UI截图共六个典型类别,共计12,095张图像。基于此,我们评估了五种代表性图像生成检测器在零样本设置下的表现,并进一步探索了多模态视觉-语言模型在此任务中的能力。结果表明,不同文本丰富领域间性能差异显著,现有检测器表现出明显的优势与失效模式。尽管最强检测器整体表现良好,但在某些结构化类别中仍无效,且对JPEG压缩高度敏感。视觉-语言模型提供有前景的补充方案,但仍难以应对高度结构化的文本丰富图像。这些发现凸显了亟需具备文本与版式感知能力的新型检测方法。数据集已公开于 https://huggingface.co/datasets/Shuyiww/TextRich。
原文摘要 · Abstract (English)
Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation models become increasingly capable of synthesizing realistic textual content and structured visual designs, detecting AI-generated text-rich images has become an important challenge for digital trust and content authenticity. Existing benchmarks, however, largely focus on object-centric images and provide limited coverage of scenarios where textual semantics and layout organization are central. In this paper, we introduce TextRich, a multi-domain benchmark for detecting text-rich images generated by OpenAI's GPT-Image-2. The benchmark contains 12,095 images across six representative categories: commercial posters, infographic charts, academic posters, receipts, tables, and UI screenshots. Using this benchmark, we evaluate five representative AI-generated image detectors under a zero-shot setting and further explore the capability of a multimodal vision-language model for this task. Our results reveal substantial performance variations across text-rich domains, where existing AI-generated image detectors exhibit distinct strengths and failure modes. Although the strongest detector achieves competitive overall performance, it remains ineffective on certain structured categories and highly sensitive to JPEG compression. Vision-language models provide a promising complementary approach, but still struggle with highly structured text-rich images. These findings highlight the need for text- and layout-aware detection methods for modern AI-generated images. Our dataset is released at https://huggingface.co/datasets/Shuyiww/TextRich.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。