arXiv:2608.00077cs.CV2026-08

提出空间溯源审计,揭示OCR关键区域剪枝后的答案可信度问题

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

论文配图:Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
图 1 · 摘自论文原文
  • 通过几何溯源追踪视觉令牌来源,检测剪枝后答案是否依赖真实文本区域
  • 30%保留率下模型准确率仅微降0.003,但有效支持覆盖率差异达2倍以上
  • 适合关注OCR任务可靠性与压缩效率平衡的研究者和工程落地人员

视觉令牌剪枝通常以固定保留率下的回答质量为评判标准。对于富含文本的多模态大语言模型(MLLM),该方法可能遗漏一种特殊失败:即使无保留令牌可追溯至支持答案的小型OCR区域,答案仍可能正确。本文将此盲点转化为基于答案行为、几何令牌起源、干预措施与实际成本的空间溯源审计。无需训练的透明选择器可定位可控操作点。在图像不重叠验证条件下,Qwen在30%保留率下的准确率为0.786,全量输入为0.783(配对图像簇差值+0.003,95%置信区间[-0.014, +0.020]),但相同预算下,目标剪枝、随机剪枝与网格剪枝的有效支持覆盖率分别为0.620、0.270和0.318。在Qwen3-VL-8B、LLaVA-1.5-7B和InternVL3.5-8B上,匹配控制、干预测试、检测器实验与外部方法揭示了模型特有的质量-风险-可追溯性边界,这些边界无法仅由准确率暴露。生成前缀实现最高4.32倍批处理预填充加速与76.4%更低的增量峰值内存;全验证TextVQA与DocVQA进一步表明,有利的目标验证点并不意味着任务通用压缩。因此,视觉令牌剪枝应同时报告剩余空间溯源性与实际成本,而不仅是质量和压缩比。

原文摘要 · Abstract (English)

Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

视觉剪枝OCR可靠性空间溯源MLLM优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。