arXiv:2602.00122cs.CVcs.AI2026-02

首个评估中英双语密集文档图像编辑能力的基准测试

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

  • 构建包含942个样本的中英双语文档编辑数据集
  • 提出基于OCR解析的细粒度评估框架,准确率提升显著
  • 适合研究多语言文档生成与编辑的开发者使用

近年来,图像编辑模型在自然语言指令下灵活操控视觉内容方面取得显著进展。然而,密集视觉文档编辑——即在保持原文本风格与背景上下文的前提下修改图像中的文字内容——仍是一个未被充分探索的方向。现有方法主要聚焦于英文场景和文本稀疏的图像,难以应对包含中英文的密集复杂文档或非拉丁文字体。为此,我们提出VDE Bench(Visual Doc Edit Bench),一个经过严格人工标注与评估的基准测试,专门用于评估图像编辑模型在中英双语及复杂视觉文档编辑任务中的表现。该基准包含942个基于指令的图像编辑样本,原始图像涵盖学术论文、海报、演示文稿、考试材料和报纸等密集中英文文本文档。此外,我们引入一种新型评估框架,系统性地在OCR解析层面量化编辑性能,实现细粒度的文字修改准确性评估。基于此基准,我们对代表性图像编辑模型进行了全面评估,人工验证表明人类判断与自动指标高度一致。VDE Bench是首个系统性评估图像编辑模型在双语密集文本视觉文档上表现的基准。

原文摘要 · Abstract (English)

In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language instructions. However, an important yet underexplored research direction remains dense visual document image editing, which involves modifying textual content within images while faithfully preserving the original text style and background context. Existing methods primarily focus on English scenarios and images with relatively sparse text, and thus cannot adequately address dense, structurally complex documents or non-Latin scripts such as Chinese. To bridge this gap, we propose VDE Bench (Visual Doc Edit Bench), a rigorously human annotated and evaluated benchmark specifically designed to assess the performance of image editing models on bilingual Chinese-English and complex visual document editing tasks. The benchmark comprises a high quality dataset of 942 instruction based image editing samples, whose seed images encompass dense Chinese and English text documents including academic papers, posters, presentation slides, examination materials, and newspapers. Furthermore, we introduce a novel evaluation framework that systematically quantifies editing performance at the OCR parsing level, thereby enabling fine grained assessment of text modification accuracy. Based on this benchmark, we conduct a comprehensive evaluation of representative image editing models. Human verification demonstrates a high degree of consistency between human judgments and automated evaluation metrics. VDE Bench constitutes the first systematic benchmark for evaluating the performance of image editing models on bilingual dense text visual documents.

文档编辑多语言OCR评估图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。