arXiv:2511.11759cs.CRcs.AI2025-11被引 3

AI模型可被破解用于诈骗老年人,实测11%成功得手。

Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation

  • 用攻击手段破解大模型安全防护,生成钓鱼内容
  • 对108名老人测试,11%受骗成功
  • 揭示当前防护体系对弱势群体形同虚设

我们展示了攻击者如何利用大模型安全漏洞伤害易受骗人群:从破解大模型生成钓鱼内容,到真实投放,最终成功欺骗老年人。系统评估了六款前沿大模型在四类攻击下的安全防线,发现多个模型对特定攻击几乎完全无防御能力。在包含108名老年志愿者的人类验证实验中,由AI生成的钓鱼邮件成功使11%的参与者受骗。本研究首次完整呈现针对老年人的攻击全流程,表明当前人工智能安全措施无法有效保护最脆弱群体。除生成钓鱼内容外,大模型还能突破语言障碍,大规模开展多轮信任构建对话,从根本上改变诈骗经济模式。尽管部分厂商报告有自愿反滥用措施,我们认为仍远远不足。

原文摘要 · Abstract (English)

We present an end-to-end demonstration of how attackers can exploit AI safety failures to harm vulnerable populations: from jailbreaking LLMs to generate phishing content, to deploying those messages against real targets, to successfully compromising elderly victims. We systematically evaluated safety guardrails across six frontier LLMs spanning four attack categories, revealing critical failures where several models exhibited near-complete susceptibility to certain attack vectors. In a human validation study with 108 senior volunteers, AI-generated phishing emails successfully compromised 11\% of participants. Our work uniquely demonstrates the complete attack pipeline targeting elderly populations, highlighting that current AI safety measures fail to protect those most vulnerable to fraud. Beyond generating phishing content, LLMs enable attackers to overcome language barriers and conduct multi-turn trust-building conversations at scale, fundamentally transforming fraud economics. While some providers report voluntary counter-abuse efforts, we argue these remain insufficient.

AI安全钓鱼攻击老年人防护大模型漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。