构建首个用于研究AR+大模型社交工程的多模态数据集
SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors
- 采集60人参与的180段模拟攻击对话,含视觉音频与环境信息
- 实测显示93.3%点击钓鱼链接,85%接受可疑电话
- 适合安全研究者、防御系统开发者和人机交互学者使用
SEAR数据集是一个新型多模态资源,用于研究通过增强现实(AR)和多模态大语言模型(LLMs)驱动的社交工程(SE)攻击。该数据集记录了60名参与者在模拟对抗场景(如会议、课堂、社交活动)中的180段标注对话,包含同步采集的AR视觉/音频线索(如面部表情、语调)、环境上下文及定制的社交媒体资料,并附有主观评估指标,如信任评分和易受攻击度。关键发现表明,该数据集揭示了令人担忧的攻击效力:93.3%的钓鱼链接点击率,85%的来电接受率,以及76.7%的互动后信任度上升。该数据集支持检测AR驱动的社交工程攻击、设计防御框架,并理解多模态对抗操纵机制。采用严格的伦理保障措施,包括匿名化处理和IRB合规,确保负责任使用。SEAR数据集已公开于https://github.com/INSLabCN/SEAR-Dataset。
原文摘要 · Abstract (English)
The SEAR Dataset is a novel multimodal resource designed to study the emerging threat of social engineering (SE) attacks orchestrated through augmented reality (AR) and multimodal large language models (LLMs). This dataset captures 180 annotated conversations across 60 participants in simulated adversarial scenarios, including meetings, classes and networking events. It comprises synchronized AR-captured visual/audio cues (e.g., facial expressions, vocal tones), environmental context, and curated social media profiles, alongside subjective metrics such as trust ratings and susceptibility assessments. Key findings reveal SEAR's alarming efficacy in eliciting compliance (e.g., 93.3% phishing link clicks, 85% call acceptance) and hijacking trust (76.7% post-interaction trust surge). The dataset supports research in detecting AR-driven SE attacks, designing defensive frameworks, and understanding multimodal adversarial manipulation. Rigorous ethical safeguards, including anonymization and IRB compliance, ensure responsible use. The SEAR dataset is available at https://github.com/INSLabCN/SEAR-Dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。