实证研究指出,去中心化AI代理的信任协议存在身份造假与声誉可操纵问题。
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

- 通过链上数据验证身份与声誉注册,发现多数身份为无效占位符
- 声誉值不可比、反馈难验证,且90%以上评分来自恶意合谋账号
- 建议重构协议机制,适合关注AI代理经济安全的研究者参考
随着自主AI代理在组织间交易日益频繁,信任问题凸显:如何评估未知对手的可信度?ERC-8004协议首次构建了无需许可的AI代理经济信任层,基于身份、声誉和验证三个链上注册表。尽管快速普及,该协议尚未有实证研究。本文对Ethereum、BNB Smart Chain(BSC)和Base三条链上的ERC-8004进行首次实证分析,覆盖从部署至2026年5月13日的数据,采集链上身份与声誉事件、离线文件及x402支付交易。结果显示,仅3%、4%、15%的身份在以太坊、BSC和Base上暴露有效ERC-8004注册文件并具备活跃服务端点;声誉注册无法作为可信信号——评分不具可比性,反馈极少基于可验证交互,且可低成本操控。一致地,73.5%、59.2%、90.6%的评审者表现出协同的Sybil行为。剔除可疑反馈后,分别有15.8%、77.9%、86.8%的被评代理无有效评分。研究提出具体改进建议,为未来协议迭代提供依据,并建立AI代理市场研究的实证基准。
原文摘要 · Abstract (English)
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknown counterpart is trustworthy? The ERC-8004 protocol addresses this challenge with the first permissionless trust layer for AI agent economies, built around three on-chain registries for Identity, Reputation, and Validation. Despite its rapid adoption, the protocol has not been studied empirically, leaving it unclear whether the information it records provides a trustworthy basis for decision-making. To address this gap, we present the first empirical study of ERC-8004 across three chains: Ethereum, BNB Smart Chain (BSC), and Base, covering the period from protocol deployment through May 13, 2026. We crawl on-chain Identity and Reputation events, off-chain files, and x402 payment transactions. On the identity side, we find that most registrations are placeholders rather than active agents, with only a small fraction (3%, 4%, and 15% across Ethereum, BSC, and Base) exposing a valid ERC-8004 registration file with at least one live service endpoint. On the reputation side, we show that the Registry, as currently deployed, cannot function as a trust signal: values are not commensurable, feedback records are rarely grounded in verifiable interactions, and reputation can be manipulated at minimal cost. Consistent with these design weaknesses, we find that a substantial fraction of reviewers (73.5%, 59.2%, and 90.6% across Ethereum, BSC, and Base) exhibit coordinated Sybil behavior. After removing Sybil-flagged feedback, 15.8%, 77.9%, and 86.8% of rated agents, respectively, are left with no valid feedback. We then turn these findings into concrete recommendations for future revisions of ERC-8004. Our study yields actionable protocol-design implications and establishes an empirical baseline for research on AI agent markets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。