用AI代理自动发现并验证智能合约漏洞,成功率63%且单个漏洞可获利超850万美元。
AI Agent Smart Contract Exploit Generation
- 构建AI代理系统A1,通过六种专用工具自主挖掘漏洞并执行验证。
- 在36个真实合约上成功生成63%的可盈利漏洞,最高单次获利859万美元。
- 揭示攻击者只需6000美元即可获利,而防御需6万美元,凸显攻防失衡。
智能合约漏洞已导致数十亿美元损失,但发现可行攻击仍具挑战性。传统模糊测试依赖僵化规则,难以应对复杂攻击;人工审计虽全面但效率低且无法扩展。大语言模型(LLM)提供了兼具人类推理与机器速度的潜在解决方案。早期研究显示,仅靠提示词生成的漏洞推测存在高误报率且未经验证。为此,我们提出A1,一个将任意LLM转化为端到端漏洞生成器的智能体系统。A1赋予代理六种领域专用工具,实现从理解合约行为到在真实区块链状态上测试策略的全流程自治。所有输出均通过执行验证,确保仅报告具有利润的原型攻击。我们在以太坊和币安智能链的36个真实漏洞合约上评估A1,于VERITE基准上取得63%的成功率。所有成功案例中,单个漏洞最高提取收益达859万美元,总计933万美元。基于历史攻击的蒙特卡洛分析表明,即时检测成功率可达86-89%,延迟一周后降至6-21%。经济分析揭示显著不对称:攻击者在6000美元级漏洞即能获利,而防御方需6万美元投入才具可行性,引发对AI代理是否必然偏向攻击的根本质疑。
原文摘要 · Abstract (English)
Smart contract vulnerabilities have led to billions in losses, yet finding actionable exploits remains challenging. Traditional fuzzers rely on rigid heuristics and struggle with complex attacks, while human auditors are thorough but slow and don't scale. Large Language Models offer a promising middle ground, combining human-like reasoning with machine speed. Early studies show that simply prompting LLMs generates unverified vulnerability speculations with high false positive rates. To address this, we present A1, an agentic system that transforms any LLM into an end-to-end exploit generator. A1 provides agents with six domain-specific tools for autonomous vulnerability discovery, from understanding contract behavior to testing strategies on real blockchain states. All outputs are concretely validated through execution, ensuring only profitable proof-of-concept exploits are reported. We evaluate A1 across 36 real-world vulnerable contracts on Ethereum and Binance Smart Chain. A1 achieves a 63% success rate on the VERITE benchmark. Across all successful cases, A1 extracts up to \$8.59 million per exploit and \$9.33 million total. Using Monte Carlo analysis of historical attacks, we demonstrate that immediate vulnerability detection yields 86-89% success probability, dropping to 6-21% with week-long delays. Our economic analysis reveals a troubling asymmetry: attackers achieve profitability at \$6,000 exploit values while defenders require \$60,000 -- raising fundamental questions about whether AI agents inevitably favor exploitation over defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。