用神经符号框架检测增强现实中的认知攻击,让安全机制既智能又可解释。
A Neurosymbolic Framework for Interpretable Cognitive Attack Detection in Augmented Reality
- 融合视觉语言模型与符号推理,构建感知图捕捉物体关系和时间上下文
- 通过粒子滤波统计推理发现语义动态异常,准确识别认知攻击
- 兼顾模型适应性与结果可解释性,适合安全敏感的AR应用
增强现实(AR)通过将虚拟元素叠加到真实世界来扩展人类感知。然而,虚拟与现实内容的紧密耦合使AR易受认知攻击:即扭曲用户对环境语义理解的操纵。现有检测方法主要关注像素或图像层面的视觉不一致,缺乏语义推理能力与可解释性。为此,我们提出CADAR,一种用于AR认知攻击检测的神经符号框架,整合了神经与符号推理。CADAR将预训练模型的多模态视觉语言表征融合为感知图,捕捉物体、关系及时间上下文显著性。在此结构基础上,基于粒子滤波的统计推理模块推断语义动态中的异常,以揭示认知攻击。该组合兼具现代视觉语言模型的适应性与概率符号推理的可解释性。在自建的AR认知攻击数据集上的初步实验表明,其性能持续优于现有方法,凸显了神经符号方法在鲁棒且可解释的AR安全中的潜力。
原文摘要 · Abstract (English)
Augmented Reality (AR) enriches human perception by overlaying virtual elements onto the physical world. However, this tight coupling between virtual and real content makes AR vulnerable to cognitive attacks: manipulations that distort users' semantic understanding of the environment. Existing detection methods largely focus on visual inconsistencies at the pixel or image level, offering limited semantic reasoning or interpretability. To address these limitations, we introduce CADAR, a neuro-symbolic framework for cognitive attack detection in AR that integrates neural and symbolic reasoning. CADAR fuses multimodal vision-language representations from pre-trained models into a perception graph that captures objects, relations, and temporal contextual salience. Building on this structure, a particle-filter-based statistical reasoning module infers anomalies in semantic dynamics to reveal cognitive attacks. This combination provides both the adaptability of modern vision-language models and the interpretability of probabilistic symbolic reasoning. Preliminary experiments on an AR cognitive-attack dataset demonstrate consistent advantages over existing approaches, highlighting the potential of neuro-symbolic methods for robust and interpretable AR security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。