提出新评估框架,精准衡量视觉文本压缩的保真度。
ZeroSense:How Vision matters in Long Context Compression
- 构建解耦评估框架,分离压缩质量与下游模型语义推理
- 设计零相关性基准测试集,避免上下文依赖干扰评估结果
- 发现压缩质量与任务准确率严重不一致,凸显评估革新必要性
近期基于视觉的文本压缩(VTC)方法,如 DeepSeek-OCR,通过文本转图像渲染,在长上下文建模任务中实现了极高的令牌压缩比。然而,现有评估协议主要依赖下游任务性能,由于多模态大模型(MLLMs)固有的语言先验,难以真实反映文本保留质量。本文提出一种新评估框架,解耦 MLLMs 的能力以精准评估 VTC 质量。在此框架下,引入 ZeroSense 基准,确保测试样本间低语义相关性。通过消除上下文依赖,该基准保证评估结果纯粹反映 VTC 性能,不受下游模型语义推断影响。跨多个数据集的大量实验表明,VTC 质量与下游任务准确率存在显著差异,凸显解耦评估框架的必要性。
原文摘要 · Abstract (English)
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs' capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating contextual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。