arXiv:2507.23226cs.CV2025-07

用视觉语言模型检测干扰用户感知的增强现实内容

Toward Safe, Trustworthy and Realistic Augmented Reality User Experience

  • 基于视觉语言模型与多模态推理,识别有害虚拟内容
  • 提出双系统可检测视觉误导与信息遮蔽攻击
  • 适合关注AR安全与用户体验的研究者和开发者

随着增强现实(AR)日益融入日常生活,确保其虚拟内容的安全性与可信度至关重要。本研究针对可能妨碍关键信息或隐性操控用户感知的任务有害型AR内容,开发了两个系统:ViDDAR与VIM-Sense,利用视觉语言模型(VLMs)和多模态推理模块实现检测。在此基础上,提出三个未来方向:虚拟内容的自动化、感知对齐的质量评估;多模态攻击检测;以及将VLMs适配于AR设备上的高效、以用户为中心部署。整体工作旨在建立可扩展、以人为本的AR体验保护框架,并寻求关于感知建模、多模态内容实现及轻量化模型适配的反馈。

原文摘要 · Abstract (English)

As augmented reality (AR) becomes increasingly integrated into everyday life, ensuring the safety and trustworthiness of its virtual content is critical. Our research addresses the risks of task-detrimental AR content, particularly that which obstructs critical information or subtly manipulates user perception. We developed two systems, ViDDAR and VIM-Sense, to detect such attacks using vision-language models (VLMs) and multimodal reasoning modules. Building on this foundation, we propose three future directions: automated, perceptually aligned quality assessment of virtual content; detection of multimodal attacks; and adaptation of VLMs for efficient and user-centered deployment on AR devices. Overall, our work aims to establish a scalable, human-aligned framework for safeguarding AR experiences and seeks feedback on perceptual modeling, multimodal AR content implementation, and lightweight model adaptation.

增强现实安全检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。