arXiv:2412.01077cs.ITcs.LG2024-12被引 12

用强化学习优化感知与通信一体化系统,提升性能。

A Memory-Based Reinforcement Learning Approach to Integrated Sensing and Communication

  • 基于深度强化学习,通过奖励机制平衡通信与感知
  • 考虑有记忆信道时,性能显著优于无记忆简化模型
  • 适合研究智能无线系统、动态资源调度的学者

本文研究点对点集成感知与通信(ISAC)系统,发射端在具有记忆特性的信道中传输信息的同时,利用回波信号实时估计信道状态。基于马塞的定向信息理论,我们建立了在线感知下的容量-失真权衡模型。由于波形优化涉及复杂联合目标,提出一种深度强化学习方法,使智能体通过学习通信增益与感知损失之差来优化性能。因状态空间理论上无界,采用深度确定性策略梯度算法(DDPG)。数值结果表明,在无界状态空间下性能显著优于受限状态空间模型;当状态空间退化为无记忆时,仅能采用无记忆波形策略。研究凸显了充分挖掘ISAC系统内在记忆特性的重要性。

原文摘要 · Abstract (English)

In this paper, we consider a point-to-point integrated sensing and communication (ISAC) system, where a transmitter conveys a message to a receiver over a channel with memory and simultaneously estimates the state of the channel through the backscattered signals from the emitted waveform. Using Massey's concept of directed information for channels with memory, we formulate the capacity-distortion tradeoff for the ISAC problem when sensing is performed in an online fashion. Optimizing the transmit waveform for this system to simultaneously achieve good communication and sensing performance is a complicated task, and thus we propose a deep reinforcement learning (RL) approach to find a solution. The proposed approach enables the agent to optimize the ISAC performance by learning a reward that reflects the difference between the communication gain and the sensing loss. Since the state-space in our RL model is à priori unbounded, we employ deep deterministic policy gradient algorithm (DDPG). Our numerical results suggest a significant performance improvement when one considers unbounded state-space as opposed to a simpler RL problem with reduced state-space. In the extreme case of degenerate state-space only memoryless signaling strategies are possible. Our results thus emphasize the necessity of well exploiting the memory inherent in ISAC systems.

ISAC强化学习通信感知一体化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。