提出CT-PPO算法,高效统筹多租户雷达通信网络的感知任务
Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC Networks
- 采用共享工作负载轨迹的因子化策略,分离奖励与约束信用
- 相比基线提升1.847平均回报,降低6.277资源成本
- 无需重训练,在高并发场景下仍保持显著优势
集成感知与通信(ISAC)网络可通过共享感知会话服务兼容请求,但合并会话涉及准入、复用、配置选择、感知服务等级协议(SLA)、通信服务质量(QoS)及未来承诺等多重耦合。本文将该问题建模为约束马尔可夫决策过程,提出共迹因子化约束近端策略优化(CT-PPO)。训练时,随机策略副本共享相同基础工作负载轨迹;留一法折扣蒙特卡洛回报对比为适用的动作因子提供奖励信用,而约束信用保持因子/前缀专属。在五组训练种子与匹配工作负载下,CT-PPO实现最高均值宏观回报,较联合信用PPO(JC-PPO)高出0.934(95%置信区间[0.702, 1.164]),较SLA感知贪心算法高出1.847;相较于JC-PPO,其感知资源成本降低6.277,每创建会话接受请求数提升0.0806。四向消融实验表明,仅因子化代理无显著回报增益,而引入共迹奖励信用带来主导提升。部署中使用公开观测与硬掩码;CT-PPO额外参数仅存在于训练端,执行器开销与JC-PPO相当,仅执行器CPU延迟基本不变。
原文摘要 · Abstract (English)
Integrated sensing and communication (ISAC) networks can serve compatible requests through shared sensing sessions, but consolidation couples admission, reuse, profile selection, sensing service-level agreements (SLAs), communication quality of service (QoS), and future commitments. We formulate this problem as a constrained Markov decision process and propose Common-Trace Factorized Constrained Proximal Policy Optimization (CT-PPO). During training, stochastic policy replicas share the same primitive workload trace; leave-one-out discounted Monte Carlo return contrasts provide reward credit to applicable actor factors, while constraint credit remains factor/prefix-specific. Across five training seeds and matched workloads, CT-PPO achieves the highest mean macro return, exceeding matched Joint-Credit PPO (JC-PPO) by 0.934 (95% confidence interval [0.702, 1.164]) and SLA-Aware Greedy by 1.847; versus JC-PPO, it reduces sensing-resource cost by 6.277 and raises accepted requests per created session by 0.0806. A four-way ablation shows that the factorized surrogate alone yields no detectable macro-return gain, whereas adding common-trace reward credit produces the dominant improvement. Without retraining, CT-PPO retains a return advantage at low, nominal, and high arrival loads, with the strongest gain under clustered arrivals. Deployment uses public observations and hard masks; CT-PPO's extra parameters are training-side, its actor footprint matches JC-PPO, and actor-only CPU latency is effectively unchanged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。