提出空间溯源审计,揭示OCR关键区域剪枝后的答案可信度问题
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

- 通过几何溯源追踪视觉令牌来源,检测剪枝后答案是否依赖真实文本区域
- 30%保留率下模型准确率仅微降0.003,但有效支持覆盖率差异达2倍以上
- 适合关注OCR任务可靠性与压缩效率平衡的研究者和工程落地人员
视觉令牌剪枝通常以固定保留率下的回答质量为评判标准。对于富含文本的多模态大语言模型(MLLM),该方法可能遗漏一种特殊失败:即使无保留令牌可追溯至支持答案的小型OCR区域,答案仍可能正确。本文将此盲点转化为基于答案行为、几何令牌起源、干预措施与实际成本的空间溯源审计。无需训练的透明选择器可定位可控操作点。在图像不重叠验证条件下,Qwen在30%保留率下的准确率为0.786,全量输入为0.783(配对图像簇差值+0.003,95%置信区间[-0.014, +0.020]),但相同预算下,目标剪枝、随机剪枝与网格剪枝的有效支持覆盖率分别为0.620、0.270和0.318。在Qwen3-VL-8B、LLaVA-1.5-7B和InternVL3.5-8B上,匹配控制、干预测试、检测器实验与外部方法揭示了模型特有的质量-风险-可追溯性边界,这些边界无法仅由准确率暴露。生成前缀实现最高4.32倍批处理预填充加速与76.4%更低的增量峰值内存;全验证TextVQA与DocVQA进一步表明,有利的目标验证点并不意味着任务通用压缩。因此,视觉令牌剪枝应同时报告剩余空间溯源性与实际成本,而不仅是质量和压缩比。
原文摘要 · Abstract (English)
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。