用几何变换检测文本合理性,防止视觉支持不足的错误输出。
Geometric Risk Control for Vision-Language Model OCR
- 通过可重复的几何变换作为探针,检测文本生成是否结构合理。
- 在多个基准上降低释放文本的平均错误率和灾难性错误率。
- 适合对准确性要求高的审计类文档识别场景。
视觉语言模型(VLMs)可实现灵活的生成式OCR,但其开放解码器可能生成缺乏视觉支持却流畅的错误文本。在审计敏感记录中,此类输出代价可能高于不输出。因此,冻结或外部部署的VLM需要一个外部决策层,以判断转录是否具备足够的视觉证据可供发布。本文提出几何风险控制器(GRC),一种模型无关的控制器,将受控几何变换视为可重复的黑箱探针,筛选结构上不合理的延续,并仅释放由跨视图一致证据支持的唯一候选。该协议在可复现的固定决策规则下,实现了可量化的选择性暴露控制,包含明确的覆盖范围与查询成本。在多个冻结VLM和标准场景文本基准上的实验表明,该方法持续降低了释放输出的均值误差、尾部误差及灾难性错误,同时保持高覆盖率。
原文摘要 · Abstract (English)
Vision-language models (VLMs) enable flexible generative optical character recognition (OCR), while their open-ended decoders can expose wrong but fluent text with weak visual support. In audit-sensitive records, such an output can be more costly than abstention. Frozen or externally served VLMs therefore require an external decision layer that can determine whether a transcription has sufficient visual evidence for release. We introduce the Geometric Risk Controller (GRC), a model-agnostic controller that treats controlled geometric transformations as repeatable black-box probes, screens structurally implausible continuations, and releases the unique candidate supported by coherent cross-view evidence. The protocol provides empirical selective exposure control with explicit coverage and query cost under a reproducible fixed decision rule. Experiments across frozen VLMs and standard scene-text benchmarks consistently reduce mean, upper-tail, and catastrophic error among released outputs while retaining high coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。