用表格强化学习优化地铁扩展,省时省碳还更公平
Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning

- 将地铁扩展问题转为非马尔可夫奖励过程,用表格强化学习求解
- 训练次数减少18倍,碳排放降低12倍,性能媲美深度强化学习
- 兼顾效率与公平性,适合城市规划与交通政策制定者
我们研究地铁网络扩展问题(MNEP),这是交通网络设计问题(TNDP)的一个子集,旨在满足出行需求。传统方法依赖精确或启发式算法,并需专家设定约束以缩小搜索空间。近年来,深度强化学习(Deep RL)因在复杂序列决策中的有效性而兴起,但其计算成本高、环境负担重,且难以解释。我们发现MNEP问题规模较小,无需深度强化学习。通过将MNEP重构为非马尔可夫奖励决策过程(NMRDP),采用表格强化学习,在显著减少训练轮次的同时实现相近性能,并提升可解释性。同时,将社会公平性指标纳入奖励函数,兼顾效率与公平。在西安和阿姆斯特丹的真实场景中评估,平均减少训练轮次18倍,碳排放降低12倍,性能仍可与深度强化学习比肩。该方法具有可复现性、模块化、可解释性和资源高效性,适用于其他组合优化问题。
原文摘要 · Abstract (English)
We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro systems to satisfy travel demand. Traditional methods rely on exact and heuristic approaches that require expert-defined constraints to reduce the search space. Recently, deep reinforcement learning (Deep RL) has emerged due to its effectiveness in complex sequential decision-making processes-it remains, however, computationally expensive, environmentally costly, and requires additional engineering to interpret. We show that MNEP problems are small enough to not require Deep RL methods. Reformulating the MNEP as a Non-Markovian Rewards Decision Process (NMRDP), we use tabular RL to achieve similar performance with significantly fewer training episodes, additionally offering greater interpretability. Additionally, we incorporate social equity criteria into the reward functions, focusing on efficiency and fairness, highlighting the versatility of our method. Evaluated in real-world settings-Xi'an and Amsterdam-our method reduces total episodes by a factor of 18 and total carbon emissions by a factor of 12 on average, while remaining competitive with Deep RL. This approach offers a replicable, modular, interpretable, and resource-efficient solution with potential applications to other combinatorial optimization problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。