arXiv:2601.04752cs.CV2026-01中稿 · ITC-CSCC 2025

用骨架化方法生成对抗扰动,挑战大模型数学文本识别能力

Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition

  • 通过骨架化压缩视觉搜索空间,精准攻击含数学公式图像
  • 在ChatGPT上成功诱导字符与语义错误,验证攻击有效性
  • 揭示大模型对复杂数学文本的视觉理解缺陷,适合安全评估者参考

本研究通过引入基于骨架化的新型对抗攻击方法,探索基础模型的视觉能力与局限性。该方法针对包含文本的图像,尤其是因LaTeX转换和复杂结构而更具挑战性的数学公式图像,有效缩小了扰动搜索空间。我们对原始图像与对抗扰动后输出之间的字符级和语义级变化进行了详细评估,深入揭示模型在视觉解读与推理方面的能力。该方法在ChatGPT上的应用进一步证明了其在真实场景中的实际影响。

原文摘要 · Abstract (English)

This work explores the visual capabilities and limitations of foundation models by introducing a novel adversarial attack method utilizing skeletonization to reduce the search space effectively. Our approach specifically targets images containing text, particularly mathematical formula images, which are more challenging due to their LaTeX conversion and intricate structure. We conduct a detailed evaluation of both character and semantic changes between original and adversarially perturbed outputs to provide insights into the models' visual interpretation and reasoning abilities. The effectiveness of our method is further demonstrated through its application to ChatGPT, which shows its practical implications in real-world scenarios.

对抗攻击视觉语言模型数学文本识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。