测试AI Agent在以太坊智能合约中发现、修复和利用漏洞的能力。
EVMbench: Evaluating AI Agents on Smart Contract Security
- 基于真实区块链环境,用自动化测试评估AI行为
- 117个真实漏洞,能端到端发现并攻击运行中的合约
- 适合安全研究者和区块链开发者参考
公共区块链上的智能合约如今管理着巨额价值,其漏洞可能导致重大损失。随着AI Agent在读写和执行代码方面能力增强,人们自然关注它们在该领域中的表现——既能提升安全性,也可能带来新风险。我们提出EVMbench,一个评估AI Agent检测、修补和利用智能合约漏洞能力的基准。该评测基于40个仓库中的117个精心挑选的漏洞,在最贴近现实的环境中,采用程序化评分方式,依据测试结果与本地以太坊执行环境下的区块链状态进行判断。我们评估了多种前沿Agent,发现它们能够对运行中的区块链实例实现从漏洞发现到攻击的全流程操作。我们开源了代码、任务和工具链,以支持持续评估相关能力,并推动未来安全研究。
原文摘要 · Abstract (English)
Smart contracts on public blockchains now manage large amounts of value, and vulnerabilities in these systems can lead to substantial losses. As AI agents become more capable at reading, writing, and running code, it is natural to ask how well they can already navigate this landscape, both in ways that improve security and in ways that might increase risk. We introduce EVMbench, an evaluation that measures the ability of agents to detect, patch, and exploit smart contract vulnerabilities. EVMbench draws on 117 curated vulnerabilities from 40 repositories and, in the most realistic setting, uses programmatic grading based on tests and blockchain state under a local Ethereum execution environment. We evaluate a range of frontier agents and find that they are capable of discovering and exploiting vulnerabilities end-to-end against live blockchain instances. We release code, tasks, and tooling to support continued measurement of these capabilities and future work on security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。