arXiv:2502.17307cs.LGcs.GT2025-02IJCAI综述被引 2

用强化学习分析区块链挖矿攻击,找安全阈值

Survey on Strategic Mining in Blockchain: A Reinforcement Learning Approach

  • 用强化学习替代传统马尔可夫模型,动态优化攻击策略
  • 能估算出发动攻击所需的最低算力门槛(如1/3)
  • 适合研究去中心化系统安全与AI结合的学者

战略性挖矿攻击(如自私挖矿)通过偏离诚实行为来最大化收益,利用区块链共识协议的漏洞。传统的马尔可夫决策过程(MDP)在现代数字经济中面临可扩展性挑战。为克服这一局限,强化学习(RL)提供了一种可扩展的替代方案,可在复杂动态环境中实现策略自适应优化。本文综述了强化学习在战略挖矿分析中的作用,对比其与基于MDP的方法。首先回顾基础MDP模型及其局限性,随后探讨能够跨多种协议学习近似最优策略的强化学习框架。在此基础上,比较不同强化学习技术在推导安全阈值(如盈利攻击所需最低攻击者算力)方面的有效性。进一步,我们对共识协议进行分类,并提出开放挑战,如多智能体动态与真实世界验证。本综述强调强化学习在应对自私挖矿问题上的潜力,涵盖协议设计、威胁检测与安全分析,并为去中心化系统与人工智能驱动分析的研究者提供战略路线图。

原文摘要 · Abstract (English)

Strategic mining attacks, such as selfish mining, exploit blockchain consensus protocols by deviating from honest behavior to maximize rewards. Markov Decision Process (MDP) analysis faces scalability challenges in modern digital economics, including blockchain. To address these limitations, reinforcement learning (RL) provides a scalable alternative, enabling adaptive strategy optimization in complex dynamic environments. In this survey, we examine RL's role in strategic mining analysis, comparing it to MDP-based approaches. We begin by reviewing foundational MDP models and their limitations, before exploring RL frameworks that can learn near-optimal strategies across various protocols. Building on this analysis, we compare RL techniques and their effectiveness in deriving security thresholds, such as the minimum attacker power required for profitable attacks. Expanding the discussion further, we classify consensus protocols and propose open challenges, such as multi-agent dynamics and real-world validation. This survey highlights the potential of reinforcement learning (RL) to address the challenges of selfish mining, including protocol design, threat detection, and security analysis, while offering a strategic roadmap for researchers in decentralized systems and AI-driven analytics.

区块链安全强化学习挖矿攻击共识协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。