为6G车联网设计信息价值模型,实现通信与控制的协同优化
Learning Value of Information towards Joint Communication and Control in 6G V2X
- 构建序贯随机决策模型(SSDP),量化信息对决策的价值
- 提出基于信息价值的通信时机、内容与方式优化框架
- 适用于自动驾驶等需要实时决策的网络化控制系统
随着蜂窝车联网(C-V2X)向第六代(6G)网络演进,联网自动驾驶车辆(CAVs)成为关键应用。利用数据驱动的机器学习,尤其是深度强化学习(DRL),有望显著提升车辆在不确定性环境下的控制与通信决策能力。这两类决策紧密关联,信息价值(VoI)是连接两者的桥梁。本文引入序贯随机决策过程(SSDP)模型,定义并评估VoI,展示其在优化CAV通信系统中的应用。具体地,我们形式化定义了SSDP模型,并证明马尔可夫决策过程(MDP)是其特例。SSDP的优势在于显式表达可用信息如何提升决策质量。针对当前VoI研究零散的问题,我们提出基于MDP、强化学习(RL)和最优控制理论的系统性建模框架,划分不同类型的VoI并讨论其估计方法。最后,提出一种结构化方法,利用多种VoI度量优化‘何时’‘何事’‘如何’通信。为此,构建了基于VoI相关奖励函数的SSDP模型。虽以简单跟车控制问题为例,但该方法在广泛网络化控制系统中具有联合优化随机序列控制与通信决策的潜力。
原文摘要 · Abstract (English)
As Cellular Vehicle-to-Everything (C-V2X) evolves towards future sixth-generation (6G) networks, Connected Autonomous Vehicles (CAVs) are emerging to become a key application. Leveraging data-driven Machine Learning (ML), especially Deep Reinforcement Learning (DRL), is expected to significantly enhance CAV decision-making in both vehicle control and V2X communication under uncertainty. These two decision-making processes are closely intertwined, with the value of information (VoI) acting as a crucial bridge between them. In this paper, we introduce Sequential Stochastic Decision Process (SSDP) models to define and assess VoI, demonstrating their application in optimizing communication systems for CAVs. Specifically, we formally define the SSDP model and demonstrate that the MDP model is a special case of it. The SSDP model offers a key advantage by explicitly representing the set of information that can enhance decision-making when available. Furthermore, as current research on VoI remains fragmented, we propose a systematic VoI modeling framework grounded in the MDP, Reinforcement Learning (RL) and Optimal Control theories. We define different categories of VoI and discuss their corresponding estimation methods. Finally, we present a structured approach to leverage the various VoI metrics for optimizing the ``When", ``What", and ``How" to communicate problems. For this purpose, SSDP models are formulated with VoI-associated reward functions derived from VoI-based optimization objectives. While we use a simple vehicle-following control problem to illustrate the proposed methodology, it holds significant potential to facilitate the joint optimization of stochastic, sequential control and communication decisions in a wide range of networked control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。