用视觉语言模型检测图像中文字篡改,填补了现有研究空白。
Detecting Text Manipulation in Images using Vision Language Models
- 利用开源与闭源视觉语言模型对比分析文本篡改检测效果。
- 闭源模型如GPT-4o表现更优,开源模型仍存在差距。
- 在真实场景和伪造证件文本上测试,验证实际应用潜力。
近期研究显示大型视觉语言模型(VLMs)在图像篡改检测中表现有效,但文字篡改检测仍被忽视。本文通过在不同文本篡改数据集上分析闭源与开源VLMs,填补这一知识空白。结果表明,尽管开源模型性能逐步逼近,但仍落后于GPT-4o等闭源模型。此外,我们评估了专为图像篡改设计的VLMs在文本篡改检测中的表现,发现其存在泛化能力不足的问题。实验涵盖真实场景文本与虚构身份证上的篡改,后者模拟了现实世界中高难度滥用情形。
原文摘要 · Abstract (English)
Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing closed- and open-source VLMs on different text manipulation datasets. Our results suggest that open-source models are getting closer, but still behind closed-source ones like GPT- 4o. Additionally, we benchmark image manipulation detection-specific VLMs for text manipulation detection and show that they suffer from the generalization problem. We benchmark VLMs for manipulations done on in-the-wild scene texts and on fantasy ID cards, where the latter mimic a challenging real-world misuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。