arXiv:2411.17353cs.LGq-fin.CP2024-11被引 1

用注意力强化学习解决闪电网络节点选择与资源分配难题

Joint Combinatorial Node Selection and Resource Allocations in the Lightning Network using Attention-based Reinforcement Learning

  • 基于Transformer的强化学习框架协同优化节点选择与资源分配
  • 在多种设置下表现优于基线,路由模拟更贴近真实闪电网络
  • 发现收益最大化与去中心化目标可共存,适合区块链研究者参考

闪电网络(LN)作为比特币扩容的第二层解决方案,近年来发展迅速。据最新统计,闪电网络锁定的总价值约为5亿美元。加入网络虽有盈利机会,但需解决涉及离散节点选择和连续资源分配的复杂组合问题。现有研究对资源分配作用关注不足,且缺乏真实的路由机制模拟。本文提出一种基于Transformer增强的深度强化学习框架,用于解决联合组合节点选择与资源分配(JCNSRA)问题。我们改进了现有环境,引入模块以提升路由机制真实性,缩小与实际闪电网络路由系统的差距,并确保与问题兼容。实验表明,该模型在多种设置下均优于多个基线与启发式方法。此外,通过部署智能体并监控演化图的中心性指标,我们发现个体收益最大化与网络去中心化目标之间不存在冲突,反而呈现正相关关系。

原文摘要 · Abstract (English)

The Lightning Network (LN) has emerged as a second-layer solution to Bitcoin's scalability challenges. The rise of Payment Channel Networks (PCNs) and their specific mechanisms incentivize individuals to join the network for profit-making opportunities. According to the latest statistics, the total value locked within the Lightning Network is approximately \$500 million. Meanwhile, joining the LN with the profit-making incentives presents several obstacles, as it involves solving a complex combinatorial problem that encompasses both discrete and continuous control variables related to node selection and resource allocation, respectively. Current research inadequately captures the critical role of resource allocation and lacks realistic simulations of the LN routing mechanism. In this paper, we propose a Deep Reinforcement Learning (DRL) framework, enhanced by the power of transformers, to address the Joint Combinatorial Node Selection and Resource Allocation (JCNSRA) problem. We have improved upon an existing environment by introducing modules that enhance its routing mechanism, thereby narrowing the gap with the actual LN routing system and ensuring compatibility with the JCNSRA problem. We compare our model against several baselines and heuristics, demonstrating its superior performance across various settings. Additionally, we address concerns regarding centralization in the LN by deploying our agent within the network and monitoring the centrality measures of the evolved graph. Our findings suggest not only an absence of conflict between LN's decentralization goals and individuals' revenue-maximization incentives but also a positive association between the two.

闪电网络强化学习去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。