arXiv:2602.07152cs.CRcs.AI2026-02被引 1

揭示AI模型中的隐藏后门,提出检测与防御新方法。

Trojans in Artificial Intelligence (TrojAI) Final Report

  • 通过权重分析和触发器逆向定位恶意后门
  • 发现模型中存在高比例的自然类后门现象
  • 为部署模型提供可操作的防御策略,适合安全研究者

IARPA启动的TrojAI项目旨在应对现代人工智能中一种新兴威胁:AI后门。这些后门是故意嵌入模型的恶意隐藏机制,可在特定触发条件下导致系统异常或被攻击者劫持。该项目历时多年,梳理了威胁本质,开创了基于权重分析和触发器逆向的检测方法,并识别出尚未解决的关键挑战。本报告总结了核心成果,包括检测方法性能、灵敏度及“自然”后门的普遍性。测试结果表明,部分模型存在显著后门风险。报告最后提出经验教训与推进AI安全研究的建议。

原文摘要 · Abstract (English)

The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi-year initiative helped to map out the complex nature of the threat, pioneered foundational detection methods, and identified unsolved challenges that require ongoing attention by the burgeoning AI security field. This report synthesizes the program's key findings, including methodologies for detection through weight analysis and trigger inversion, as well as approaches for mitigating Trojan risks in deployed models. Comprehensive test and evaluation results highlight detector performance, sensitivity, and the prevalence of "natural" Trojans. The report concludes with lessons learned and recommendations for advancing AI security research.

AI安全后门检测模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。