提出新评估框架,精准衡量图文压缩中文字保留质量。
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

- 分离语义与视觉能力,避免模型先验干扰评估
- 构建零语义关联测试集,确保结果仅反映压缩质量
- 揭示压缩质量与下游任务表现严重不匹配
近期的图文压缩(VTC)方法,如DeepSeek-OCR,通过文本转图像渲染,在长上下文建模任务中实现了极高的词元压缩比。然而,现有评估协议主要依赖下游任务性能,因多模态大模型(MLLMs)固有的语言先验,无法准确衡量文本保留程度。本文提出一种解耦框架,将语义理解能力与视觉编码分离,以忠实评估VTC质量。在此框架下,引入ZeroSense基准测试集,确保测试样本间语义相关性极低。通过消除文本依赖,评估结果仅反映压缩本身质量,不受下游模型语义推断能力影响。在多个数据集上的实验证明,压缩质量与下游任务准确率存在显著分歧,凸显解耦评估框架的必要性。
原文摘要 · Abstract (English)
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs' capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating textual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。