构建AR环境下对抗性虚拟内容攻击的评测基准,检验视觉语言模型鲁棒性。
Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality
- 提出ContrAR基准,模拟真实AR场景中的矛盾虚拟内容攻击
- 在312段真实AR视频上测试11个VLM,发现检测准确率仍有提升空间
- 适合关注AR安全与多模态模型抗干扰能力的研究者
增强现实(AR)在过去十年中迅速发展。随着AR日益融入日常生活,其安全性和可靠性成为关键挑战。其中,矛盾虚拟内容攻击——即恶意或不一致的虚拟元素被引入用户视野——构成独特风险,可能导致误导、语义混淆或传播有害信息。本文系统建模此类攻击,提出ContrAR,一个用于评估视觉语言模型(VLMs)在AR环境中对虚拟内容操纵和矛盾识别能力的新基准。ContrAR包含312段经10名人类参与者验证的真实AR视频。我们进一步对11个VLM(含商业与开源模型)进行了基准测试。实验结果表明,尽管当前VLM对矛盾虚拟内容具备合理理解能力,但在检测和推理对抗性内容操纵方面仍存改进空间。此外,检测准确率与延迟之间的平衡仍具挑战性。
原文摘要 · Abstract (English)
Augmented reality (AR) has rapidly expanded over the past decade. As AR becomes increasingly integrated into daily life, its security and reliability emerge as critical challenges. Among various threats, contradictory virtual content attacks, where malicious or inconsistent virtual elements are introduced into the user's view, pose a unique risk by misleading users, creating semantic confusion, or delivering harmful information. In this work, we systematically model such attacks and present ContrAR, a novel benchmark for evaluating the robustness of vision-language models (VLMs) against virtual content manipulation and contradiction in AR. ContrAR contains 312 real-world AR videos validated by 10 human participants. We further benchmark 11 VLMs, including both commercial and open-source models. Experimental results reveal that while current VLMs exhibit reasonable understanding of contradictory virtual content, room still remains for improvement in detecting and reasoning about adversarial content manipulations in AR environments. Moreover, balancing detection accuracy and latency remains challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。