让人类反馈与低延迟共存,实现智能语义通信调度
Latency-aware Human-in-the-Loop Reinforcement Learning for Semantic Communications
- 用强化学习融合人类反馈和时延控制,动态调整语义传输
- 在多用户异构时延场景下,100%满足用户时延约束
- 适合沉浸式、高安全要求的实时通信系统应用
语义通信可实现任务对齐传输,但在沉浸式和高安全场景中需兼顾语义保真度与严格时延要求。本文提出一种时间约束的人类在环强化学习框架(TC-HITL-RL),将人类反馈、语义效用与时延控制嵌入语义感知的开放无线接入网(RAN)架构中。将人类反馈驱动的语义自适应建模为带约束的马尔可夫决策过程(CMDP),状态包含语义质量、人类偏好、队列余量与信道动态,采用带动作屏蔽和时延感知奖励设计的原偶对偶近端策略优化算法求解。所获策略在保持近似PPO级语义奖励的同时,显著降低空口及近实时RAN智能控制器处理资源的波动性。点对多点链路仿真显示,该方法在异构截止时限下稳定满足用户时延约束,优于基线调度器的奖励表现并稳定资源消耗,为低时延语义自适应提供可行蓝图。
原文摘要 · Abstract (English)
Semantic communication promises task-aligned transmission but must reconcile semantic fidelity with stringent latency guarantees in immersive and safety-critical services. This paper introduces a time-constrained human-in-the-loop reinforcement learning (TC-HITL-RL) framework that embeds human feedback, semantic utility, and latency control within a semantic-aware Open radio access network (RAN) architecture. We formulate semantic adaptation driven by human feedback as a constrained Markov decision process (CMDP) whose state captures semantic quality, human preferences, queue slack, and channel dynamics, and solve it via a primal--dual proximal policy optimization algorithm with action shielding and latency-aware reward shaping. The resulting policy preserves PPO-level semantic rewards while tightening the variability of both air-interface and near-real-time RAN intelligent controller processing budgets. Simulations over point-to-multipoint links with heterogeneous deadlines show that TC-HITL-RL consistently meets per-user timing constraints, outperforms baseline schedulers in reward, and stabilizes resource consumption, providing a practical blueprint for latency-aware semantic adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。