arXiv:2504.16116cs.CRcs.AI2025-04KDD被引 2

为区块链领域设计的AI评估基准,检验大模型真实推理能力。

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain

  • 构建涵盖9个子领域的综合评测体系,模拟真实Web3场景
  • 31个模型测试显示:基础概念掌握好,安全审计等高阶任务差
  • 强调真实推理而非记忆,适合研究安全与可信AI的人参考

Web3生态系统依托密码学原语和去中心化共识,是高风险环境,软件漏洞和激励错配会直接导致财务损失。随着大语言模型(LLMs)被用于智能合约审计、去中心化金融分析等任务,其可靠性至关重要。但通用基准无法捕捉该领域所需的专门推理能力。为此,我们提出DMind Benchmark,一个全面的评估套件,用于严格测试大模型在Web3全栈中的表现。该基准包含9个不同子领域(如基础设施、智能合约、代币经济等),结合客观知识检索与复杂开放式推理任务,模拟真实运营挑战。我们对31个主流专有及开源模型进行了广泛评估,采用防污染流程,并通过跨评委一致性检查验证评分协议的统计稳健性。分析揭示关键矛盾:模型在基础基础设施概念上表现良好,但在安全审计等高阶推理任务中存在显著缺陷。此外,我们提供帕累托分析以指导成本效益部署,并通过对抗实验表明,高分需真正推理而非表面记忆。自2025年4月开源以来,DMind Benchmark在Hugging Face连续近一周位居趋势榜第一,至2026年6月下载量超1.3万次,已成为推进Web3中安全可信AI的标准。

原文摘要 · Abstract (English)

The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignments translate directly into financial loss. As Large Language Models (LLMs) are increasingly integrated into this domain for tasks ranging from smart contract auditing to decentralized finance analytics, ensuring their reliability is paramount. However, general-purpose benchmarks fail to capture the specialized reasoning required for these adversarial and protocol-driven settings. To bridge this gap, we introduce DMind Benchmark, a comprehensive evaluation suite designed to rigorously assess LLM proficiency across the Web3 stack. DMind Benchmark encompasses nine distinct subdomains (spanning infrastructure, smart contracts, token economics, etc.) and combines objective knowledge retrieval with complex open-ended reasoning tasks that emulate real-world operational challenges. We conduct an extensive evaluation of 31 leading proprietary and open-weights models, employing a contamination-aware pipeline and verifying the statistical robustness of our scoring protocol through rigorous cross-judge consistency checks. Our analysis reveals a critical dichotomy: while models demonstrate competence in foundational infrastructure concepts, they exhibit significant vulnerabilities in high-reasoning tasks such as security auditing. Furthermore, we provide a Pareto analysis to guide cost-effective deployment and demonstrate through adversarial experiments that high performance on DMind Benchmark necessitates genuine reasoning rather than superficial memorization. Since its open-source release in April 2025, DMind Benchmark achieved the #1 trending position on Hugging Face for nearly a week and accumulated over 13k downloads by June 2026, establishing itself as a standard for advancing secure and trustworthy AI in Web3.

Web3大模型评估安全审计推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。