arXiv:2504.13209cs.CRcs.AI2025-04AAAI被引 7

用AR+多模态大模型模拟社交工程攻击,验证其高成功率与信任构建能力。

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

  • 提出SEAR框架,融合视觉听觉环境信息生成社交情境
  • 93.3%受试者在模拟中暴露于邮件钓鱼,85%愿接听攻击者来电
  • 适合安全研究者、防御系统开发者参考,警示下一代增强现实威胁

增强现实(AR)与多模态大语言模型(LLMs)的融合正重塑人机交互,但也带来新型社交工程攻击风险。本文首次系统探究利用多模态大模型驱动AR社交工程攻击的可行性,提出SEAR框架,包含三个阶段:(1) 基于AR的社会情境合成,融合视觉、听觉与环境线索;(2) 面向角色的多模态RAG,动态检索并整合上下文数据,保持角色差异性;(3) ReInteract社交工程代理,通过推理交互循环执行自适应多阶段策略。我们开展了一项经IRB批准的研究,60名参与者参与三种配置(无辅助、AR+LLM、完整SEAR流程),构建了180条标注对话的新数据集。结果表明,SEAR能有效诱导高风险行为(如93.3%参与者易受邮件钓鱼),在建立信任方面尤为显著(85%目标愿接受攻击者来电)。同时识别出“偶发生硬”等真实性不足问题。本工作为AR-LLM驱动的社交工程攻击提供了概念验证,并为应对下一代增强现实威胁提供防御思路。

原文摘要 · Abstract (English)

Augmented Reality (AR) and Multimodal Large Language Models (LLMs) are rapidly evolving, providing unprecedented capabilities for human-computer interaction. However, their integration introduces a new attack surface for social engineering. In this paper, we systematically investigate the feasibility of orchestrating AR-driven Social Engineering attacks using Multimodal LLM for the first time, via our proposed SEAR framework, which operates through three key phases: (1) AR-based social context synthesis, which fuses Multimodal inputs (visual, auditory and environmental cues); (2) role-based Multimodal RAG (Retrieval-Augmented Generation), which dynamically retrieves and integrates contextual data while preserving character differentiation; and (3) ReInteract social engineering agents, which execute adaptive multiphase attack strategies through inference interaction loops. To verify SEAR, we conducted an IRB-approved study with 60 participants in three experimental configurations (unassisted, AR+LLM, and full SEAR pipeline) compiling a new dataset of 180 annotated conversations in simulated social scenarios. Our results show that SEAR is highly effective at eliciting high-risk behaviors (e.g., 93.3% of participants susceptible to email phishing). The framework was particularly effective in building trust, with 85% of targets willing to accept an attacker's call after an interaction. Also, we identified notable limitations such as ``occasionally artificial'' due to perceived authenticity gaps. This work provides proof-of-concept for AR-LLM driven social engineering attacks and insights for developing defensive countermeasures against next-generation augmented reality threats.

社交工程AR攻击多模态大模型安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。