arXiv:2511.10291cs.ITcs.LG2025-11

用因果模型提升物联网信道接入的采样效率与可解释性

Causal Model-Based Reinforcement Learning for Sample-Efficient IoT Channel Access

  • 基于因果模型构建可解释的强化学习框架
  • 减少58%环境交互,收敛速度更快
  • 适合资源受限的物联网系统部署

尽管多智能体强化学习(MARL)在无线通信如介质访问控制(MAC)中具有优势,但在物联网(IoT)实际部署中受制于样本效率低。传统模型基于强化学习(MBRL)依赖黑箱模型,不可解释且无法推理。本文提出一种新型因果模型驱动的MARL框架,利用结构因果模型(SCMs)和注意力机制显式建模网络变量间的因果关系:控制消息如何影响观测、传输动作如何决定结果、信道观测如何影响奖励。通过学习到的因果模型进行数据增强,生成合成轨迹用于近端策略优化(PPO)策略优化。分析表明,因果MBRL相比黑箱方法实现指数级样本复杂度降低。大量仿真显示,该方法平均减少58%环境交互,收敛更快。同时,基于注意力的因果归因可揭示驱动策略的关键网络条件,提供可解释调度决策。该框架兼具高效性与可解释性,适用于资源受限的无线系统。

原文摘要 · Abstract (English)

Despite the advantages of multi-agent reinforcement learning (MARL) for wireless use case such as medium access control (MAC), their real-world deployment in Internet of Things (IoT) is hindered by their sample inefficiency. To alleviate this challenge, one can leverage model-based reinforcement learning (MBRL) solutions, however, conventional MBRL approaches rely on black-box models that are not interpretable and cannot reason. In contrast, in this paper, a novel causal model-based MARL framework is developed by leveraging tools from causal learn- ing. In particular, the proposed model can explicitly represent causal dependencies between network variables using structural causal models (SCMs) and attention-based inference networks. Interpretable causal models are then developed to capture how MAC control messages influence observations, how transmission actions determine outcomes, and how channel observations affect rewards. Data augmentation techniques are then used to generate synthetic rollouts using the learned causal model for policy optimization via proximal policy optimization (PPO). Analytical results demonstrate exponential sample complexity gains of causal MBRL over black-box approaches. Extensive simulations demonstrate that, on average, the proposed approach can reduce environment interactions by 58%, and yield faster convergence compared to model-free baselines. The proposed approach inherently is also shown to provide interpretable scheduling decisions via attention-based causal attribution, revealing which network conditions drive the policy. The resulting combination of sample efficiency and interpretability establishes causal MBRL as a practical approach for resource-constrained wireless systems.

强化学习因果模型物联网信道接入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。