用视觉语言模型检测增强现实中的有害虚拟内容,提升真实世界信息识别准确性。
ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality
- 基于视觉语言模型构建全参考检测系统,实时分析虚拟内容对任务的干扰。
- 在遮挡攻击上达到92.15%准确率,延迟仅533毫秒;信息干扰检测准确率达82.46%,延迟9.62秒。
- 首个利用VLM检测AR中任务有害内容的系统,适合安全敏感型增强现实应用。
在增强现实(AR)中,虚拟内容可提升用户体验,但不当设计或放置会损害任务表现,阻碍用户对真实世界信息的准确理解。本文研究两类任务有害虚拟内容:遮挡攻击(虚拟内容遮蔽真实物体)与信息操纵攻击(干扰真实信息解读)。我们提出数学框架表征这些攻击,并构建了开源评估数据集。为应对这些问题,提出ViDDAR(基于视觉语言模型的增强现实任务有害内容检测系统),采用用户-边缘-云架构,结合深度学习技术实现低延迟监控。据我们所知,ViDDAR是首个在AR环境中使用视觉语言模型检测任务有害内容的系统。评估结果表明,该系统能有效理解复杂场景,遮挡检测准确率达92.15%,延迟533毫秒;信息操纵内容检测准确率达82.46%,延迟9.62秒。
原文摘要 · Abstract (English)
In Augmented Reality (AR), virtual content enhances user experience by providing additional information. However, improperly positioned or designed virtual content can be detrimental to task performance, as it can impair users' ability to accurately interpret real-world information. In this paper we examine two types of task-detrimental virtual content: obstruction attacks, in which virtual content prevents users from seeing real-world objects, and information manipulation attacks, in which virtual content interferes with users' ability to accurately interpret real-world information. We provide a mathematical framework to characterize these attacks and create a custom open-source dataset for attack evaluation. To address these attacks, we introduce ViDDAR (Vision language model-based Task-Detrimental content Detector for Augmented Reality), a comprehensive full-reference system that leverages Vision Language Models (VLMs) and advanced deep learning techniques to monitor and evaluate virtual content in AR environments, employing a user-edge-cloud architecture to balance performance with low latency. To the best of our knowledge, ViDDAR is the first system to employ VLMs for detecting task-detrimental content in AR settings. Our evaluation results demonstrate that ViDDAR effectively understands complex scenes and detects task-detrimental content, achieving up to 92.15% obstruction detection accuracy with a detection latency of 533 ms, and an 82.46% information manipulation content detection accuracy with a latency of 9.62 s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。