arXiv:2504.01048cs.CVcs.CR2025-04中稿 · COLM被引 1

水印会严重干扰视觉语言模型的文档理解能力,最高导致36%性能下降。

How does Watermarking Affect Visual Language Models in Document Understanding?

  • 构建新评估框架,系统测试水印对多类文档中VLMs的影响。
  • 水印使模型性能最高下降36%,分散式水印比集中式破坏更大。
  • 水印通过扰乱注意力分布和改变语义表征影响模型,适合关注鲁棒性的研究者阅读。

视觉语言模型(VLMs)已成为金融、法律、学术等领域复杂多模态文档处理的基础模型。然而,文档常包含水印等噪声信息,引发关键问题:水印是否降低VLMs在文档理解中的表现?为此,我们提出一种新型评估框架,系统考察不同文档数据类型、水印位置及内容变化对VLMs性能的影响。实验结果表明,水印可显著损害模型表现,性能下降最高达36%。研究发现,分散式水印比集中式造成更强干扰,且含语义内容的水印比单纯视觉遮挡更具破坏性。通过注意力机制分析与嵌入相似性检验,我们确认性能下降主要源于水印导致的广泛注意力重分配以及嵌入空间中的语义表示偏移。本研究揭示了VLMs在真实文档部署中的重大挑战,并为开发抗水印推理机制提供了重要启示。

原文摘要 · Abstract (English)

Visual Language Models (VLMs) have become foundational models for document understanding tasks, widely used in the processing of complex multimodal documents across domains such as finance, law, and academia. However, documents often contain noise-like information, such as watermarks, which inevitably leads us to inquire: \emph{Do watermarks degrade the performance of VLMs in document understanding?} To address this, we propose a novel evaluation framework to investigate the effect of visible watermarks on VLMs performance. We takes into account various factors, including different types of document data, the positions of watermarks within documents and variations in watermark content. Our experimental results reveal that VLMs performance can be significantly compromised by watermarks, with performance drop rates reaching up to 36\%. We discover that \emph{scattered} watermarks cause stronger interference than centralized ones, and that \emph{semantic contents} in watermarks creates greater disruption than simple visual occlusion. Through attention mechanism analysis and embedding similarity examination, we find that the performance drops are mainly attributed to that watermarks 1) force widespread attention redistribution, and 2) alter semantic representation in the embedding space. Our research not only highlights significant challenges in deploying VLMs for document understanding, but also provides insights towards developing robust inference mechanisms on watermarked documents.

视觉语言模型文档理解水印干扰模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。