测试AI防御模型在实时社交工程攻击中的真实防护能力
Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction

- 构建300例在线租房场景,区分4类信任链失败结构
- 干预率最高达96.3%,但识别错误与行动脱节现象普遍
- 适合研究安全AI对抗实时诈骗的团队参考
生成式AI使社交工程攻击更流畅、自适应且可规模化,亟需基于大模型的实时防御机制。本文探讨防御模型是识别风险根源还是仅响应表层线索。提出信任链定位框架,判断交互失败发生在角色权限、资产控制、验证充分性或交易路径环节。构建包含20种场景、300个案例的受控在线租房语料库,涵盖合法案例、四种结构失效模式及三种表面条件。在状态化轮次交互和单次静态两种设定下,评估五种防御模型,共产生3000次评测结果。尽管无模型出现明确不安全合规,但防御效果差异显著:干预率从0%至96.3%不等。保护行为与正确结构定位常不一致,部分模型误判关键环节或识别失败却未采取行动。资产控制失效是主要定位瓶颈,各模型对表面特征敏感度不同,实时与静态设置表现差异因模型而异。结果表明,表面安全行为不足以代表有效防护;实时反诈能力需独立评估干预行为、时机、结构定位及误报率。
原文摘要 · Abstract (English)
Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions. We ask whether such defenders identify the structural source of risk or merely react to surface cues. We formalize trust-chain localization: identifying whether an interaction fails at actor authority, asset control, verification sufficiency, or transaction path. We construct a controlled 300-case online-housing corpus spanning 20 scenario families, legitimate cases, four structural failure modes, and three surface conditions. Five defender models are evaluated on the same corpus in state- ful turn-by-turn and one-shot static settings, yielding 1,500 model-case evaluations per protocol and 3,000 in total. No model produced explicit unsafe compliance, yet defensive effectiveness varied sharply: intervention rates ranged from 0% to 96.3%. Protective action and correct structural localization were frequently decoupled, with models sometimes intervening while identifying the wrong trust component or recognizing a structural failure without taking protective action. Asset-control failures were a major localization bottleneck, surface sensitivity varied across models, and live-static differences were model-dependent. These findings show that safe-looking behavior alone is insufficient; live scam resistance must separately measure intervention, timing, structural localization, and false-positive behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。