arXiv:2507.20356cs.CV2025-07中稿 · the 2025 IEEE Inte…被引 12

提出多模态方法检测增强现实中的隐性视觉欺骗,准确率达88.94%

Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach

  • 融合视觉语言模型与OCR,实现跨模态语义理解
  • 在452组视频上达到88.94%攻击检测准确率
  • 适用于移动AR场景,延迟低于7.2秒

增强现实(AR)中的虚拟内容可能引入误导或有害信息,造成语义误解或用户误判。本文聚焦于视觉信息操纵(VIM)攻击,即虚拟内容以微妙方式改变真实场景的含义。我们构建了一个分类体系,将攻击分为三类形式:角色、短语、图案操纵,以及三类目的:信息替换、信息模糊、错误信息附加。基于此,我们构建了首个公开数据集AR-VIM,包含452对原始-增强现实视频,覆盖202个真实场景。为检测此类攻击,提出多模态语义推理框架VIM-Sense,结合视觉语言模型(VLMs)与基于OCR的文本分析。在AR-VIM数据集上,该系统检测准确率达88.94%,显著优于仅视觉或仅文本的基线模型。在模拟视频处理框架中平均检测延迟为7.07秒,在真实移动端Android AR应用测试中为7.17秒。

原文摘要 · Abstract (English)

The virtual content in augmented reality (AR) can introduce misleading or harmful information, leading to semantic misunderstandings or user errors. In this work, we focus on visual information manipulation (VIM) attacks in AR, where virtual content changes the meaning of real-world scenes in subtle but impactful ways. We introduce a taxonomy that categorizes these attacks into three formats: character, phrase, and pattern manipulation, and three purposes: information replacement, information obfuscation, and extra wrong information. Based on the taxonomy, we construct a dataset, AR-VIM, which consists of 452 raw-AR video pairs spanning 202 different scenes, each simulating a real-world AR scenario. To detect the attacks in the dataset, we propose a multimodal semantic reasoning framework, VIM-Sense. It combines the language and visual understanding capabilities of vision-language models (VLMs) with optical character recognition (OCR)-based textual analysis. VIM-Sense achieves an attack detection accuracy of 88.94% on AR-VIM, consistently outperforming vision-only and text-only baselines. The system achieves an average attack detection latency of 7.07 seconds in a simulated video processing framework and 7.17 seconds in a real-world evaluation conducted on a mobile Android AR application.

增强现实视觉安全多模态对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。