arXiv:2604.06914cs.LG2026-04

通过等变强化学习,让路侧单元协同优化多模态交通数据采集。

Equivariant Multi-agent Reinforcement Learning for Multimodal Vehicle-to-Infrastructure Systems

  • 基于旋转对称性设计自监督感知框架,提取车辆局部位置。
  • 多模态数据下实现两倍精度提升,性能比基线高出50%以上。
  • 适合车联网、智能交通系统研究者,解决分布式协同难题。

本文研究一种车路协同(V2I)系统,其中分布式的基站(BS)作为路侧单元(RSU)从移动车辆中收集无线与视觉多模态数据。针对去中心化的速率最大化问题,各RSU仅依赖本地观测优化资源,同时需协同保障整体网络性能。将该问题重构为分布式多智能体强化学习(MARL)问题,并引入车辆位置的旋转对称性。为此提出一种新型自监督学习框架,使每个BS代理将其多模态观测的隐层特征对齐,以提取本地区域内的车辆位置。基于此感知数据,采用具有消息传递层的图神经网络(GNN)训练等变策略网络,使各智能体可本地计算策略,同时通过信号传输机制克服部分可观测性,保证全局策略的等变性。在结合射线追踪与计算机图形学的仿真环境中进行数值实验,结果表明,所提自监督多模态感知方法具有强泛化能力,精度超过基线两倍;等变MARL训练效率显著,性能优于标准方法50%以上。

原文摘要 · Abstract (English)

In this paper, we study a vehicle-to-infrastructure (V2I) system where distributed base stations (BSs) acting as road-side units (RSUs) collect multimodal (wireless and visual) data from moving vehicles. We consider a decentralized rate maximization problem, where each RSU relies on its local observations to optimize its resources, while all RSUs must collaborate to guarantee favorable network performance. We recast this problem as a distributed multi-agent reinforcement learning (MARL) problem, by incorporating rotation symmetries in terms of vehicles' locations. To exploit these symmetries, we propose a novel self-supervised learning framework where each BS agent aligns the latent features of its multimodal observation to extract the positions of the vehicles in its local region. Equipped with this sensing data at each RSU, we train an equivariant policy network using a graph neural network (GNN) with message passing layers, such that each agent computes its policy locally, while all agents coordinate their policies via a signaling scheme that overcomes partial observability and guarantees the equivariance of the global policy. We present numerical results carried out in a simulation environment, where ray-tracing and computer graphics are used to collect wireless and visual data. Results show the generalizability of our self-supervised and multimodal sensing approach, achieving more than two-fold accuracy gains over baselines, and the efficiency of our equivariant MARL training, attaining more than 50% performance gains over standard approaches.

多智能体车联网等变学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。