用纳什谈判优化电动车间能源交易,提升效率与公平性。
Incentive-Aligned Vehicle-to-Vehicle Energy Trading via Nash-Integrated Multi-Agent Reinforcement Learning

- 将纳什谈判引入多智能体强化学习,实现双赢定价。
- 社会福利提升61.6%,交易量增62.9%,公平性指标提高40.1%。
- 适用于大规模动态车队,定价稳定且可扩展。
车对车(V2V)能源交易使电动车之间可进行去中心化的点对点能源交换,降低电网依赖并盘活过剩储能。然而,协调具有不同充电需求和不确定到离时间的自利电动车代理仍具挑战。现有方法或需集中式优化而计算受限,或缺乏公平性保障。本文将纳什讨价还价解融入多智能体深度确定性策略梯度,提出纳什-联合多智能体强化学习(Nash-MADDPG),通过纳什谈判确定高效双边定价,并以纳什引导的价格接近奖励推动智能体学习最优谈判策略。30天连续运行评估显示,相比双拍卖机制,社会福利提升61.6%,交易量增长62.9%,公平性(如Jain指数)改善40.1%。在6至100个智能体、30天持续轮换场景下测试,验证了算法在群体规模变化下的可扩展性与价格稳定性,逼近纳什讨价还价基准。
原文摘要 · Abstract (English)
Vehicle-to-vehicle (V2V) energy trading enables decentralized peer-to-peer energy exchange among electric vehicles (EVs), reducing grid dependency while monetizing surplus capacity. However, coordinating self-interested EV agents with diverse charging needs and uncertain arrival-departure schedules remains challenging. Existing approaches either require centralized optimization with computational limitations or lack fairness guarantees. This paper integrates Nash Bargaining Solution into Multi-Agent Deep Deterministic Policy Gradient, namely Nash-MADDPG, for incentive-aligned V2V energy trading. Nash bargaining determines efficient bilateral pricing, while Nash-guided price proximity rewards align agent learning toward bargaining-optimal strategies. Evaluation over 30-day continuous operation demonstrates an improvement of 61.6% in social welfare and 62.9% improvement in trading volume over Double Auction, while achieving superior fairness, such as 40.1% improvement in Jain's index. Testing across 6-100 agents over a 30-day horizon with continuous vehicle turnover confirms scalability across population size and empirically stable pricing near the Nash Bargaining benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。