arXiv:2502.15713cs.NIcs.AI2025-02被引 18

用强化学习与区块链构建无人机车联网的可信协同框架。

UAV-assisted Internet of Vehicles: A Framework Empowered by Reinforcement Learning and Blockchain

  • 双侧偏好匹配车辆与无人机,基于QoU和QoV选中继节点。
  • 多智能体强化学习实现无人机自主协同,提升网络覆盖与连接稳定性。
  • 区块链保障交互透明可追溯,适合高动态车联网场景。

本文针对无人机辅助车联网(IoV)中的中继节点选择与协同问题,提出一种融合强化学习与区块链的综合框架。该框架包含三部分:基于无人机质量(QoU)与车辆质量(QoV)的双向选择机制,用于实现车辆与无人机间的双向偏好匹配;基于近端策略优化(PPO)的去中心化多智能体深度强化学习(MDRL)模型,用于控制无人机移动并维持网络覆盖与连通性;以及基于区块链的实现,确保车-机交互过程的透明性与可追溯性。实验表明,所提方法显著提升了所选中继节点的稳定性,并最大化了无人机的覆盖范围与网络连通性。

原文摘要 · Abstract (English)

This paper addresses the challenges of selecting relay nodes and coordinating among them in UAV-assisted Internet-of-Vehicles (IoV). The selection of UAV relay nodes in IoV employs mechanisms executed either at centralized servers or decentralized nodes, which have two main limitations: 1) the traceability of the selection mechanism execution and 2) the coordination among the selected UAVs, which is currently offered in a centralized manner and is not coupled with the relay selection. Existing UAV coordination methods often rely on optimization methods, which are not adaptable to different environment complexities, or on centralized deep reinforcement learning, which lacks scalability in multi-UAV settings. Overall, there is a need for a comprehensive framework where relay selection and coordination are coupled and executed in a transparent and trusted manner. This work proposes a framework empowered by reinforcement learning and Blockchain for UAV-assisted IoV networks. It consists of three main components: a two-sided UAV relay selection mechanism for UAV-assisted IoV, a decentralized Multi-Agent Deep Reinforcement Learning (MDRL) model for autonomous UAV coordination, and a Blockchain implementation for transparency and traceability in the interactions between vehicles and UAVs. The relay selection considers the two-sided preferences of vehicles and UAVs based on the Quality-of-UAV (QoU) and the Quality-of-Vehicle (QoV). Upon selection of relay UAVs, the decentralized coordination between them is enabled through an MDRL model trained to control their mobility and maintain the network coverage and connectivity using Proximal Policy Optimization (PPO). The evaluation results demonstrate that the proposed selection and coordination mechanisms improve the stability of the selected relays and maximize the coverage and connectivity achieved by the UAVs.

无人机车联网强化学习区块链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。