arXiv:2603.06607cs.MAcs.AI2026-03

通过基准测试揭示车联网资源分配中MARL的核心挑战与算法优劣。

Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

  • 构建分层干扰博弈框架,逐步引入马尔可夫决策难题。
  • 实测显示拓扑泛化能力差导致性能下降最高达59个百分点。
  • 对比发现演员-评论家方法在复杂场景下比值函数方法强42%。

车联网(C-V2X)中的无线资源分配(RRA)至关重要,车辆需共享有限频谱以保障安全通信。多智能体强化学习(MARL)虽具潜力,但非平稳性、协调困难、动作空间大、部分可观测性及鲁棒性不足等问题相互交织,难以独立评估其影响。现有研究多集中于新算法开发,系统性基准测试较少。为此,本文将C-V2X RRA建模为逐级引入核心挑战的多智能体干扰博弈,并基于SUMO生成的高速公路轨迹构建训练与测试数据集,涵盖多样车流拓扑与干扰条件。在此基准上,评估了涵盖值函数、演员-评论家、独立学习(IL)和集中训练分散执行(CTDE)等范式的代表性MARL算法。结果表明,跨拓扑泛化能力差是主要瓶颈,平均归一化回报下降最高达59个百分点;在最复杂任务中,最优演员-评论家方法优于最优值函数方法42%。本工作揭示不同范式优劣,开源代码、数据集与基准套件,为车载网络MARL研究提供可复现的基础。

原文摘要 · Abstract (English)

Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited wireless resources to support safety-critical communications. Multi-agent reinforcement learning (MARL) has emerged as a promising approach for this problem. However, key MARL challenges, including non-stationarity, coordination difficulty, large action space, partial observability, and limited robustness and generalization, are often intertwined, making it difficult to assess their individual impact on performance in vehicular environments. Moreover, existing studies primarily focus on developing new algorithms, while systematic benchmarking and comparative analyses remain limited. To address this gap, we formulate C-V2X RRA as a hierarchy of multi-agent interference games that progressively introduce key MARL challenges. Based on this framework, we develop a suite of benchmark learning tasks and construct training and testing datasets from SUMO-generated highway traces with diverse vehicular topologies and interference conditions. Using the proposed benchmark, we evaluate representative MARL algorithms spanning value-based, actor-critic, Independent Learning (IL), and Centralized Training with Decentralized Execution (CTDE) paradigms. The results identify robustness and generalization across diverse vehicular topologies as the dominant challenge among those considered in this work, reducing average normalized return by up to 59 percentage points, and show that, on the most challenging task, the best actor-critic method outperforms the best value-based method by 42\%. By revealing the relative strengths and limitations of different MARL paradigms and open-sourcing the code, datasets, and benchmark suite, this work provides a systematic and reproducible foundation for evaluating and advancing MARL algorithms in vehicular networks.

车联网MARL资源分配基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。