arXiv:2601.22044cs.NIcs.AI2026-01中稿 · IEEE INFOCOM 2026被引 3

让预测驱动的网络控制模型变得可解释,提升透明度与调控效率。

SIA: Symbolic Interpretability for Anticipatory Deep Reinforcement Learning in Network Control

  • 用符号推理+关键指标知识图谱实时解析预测增强型强化学习决策
  • 实现亚毫秒级解释速度,比现有方法快200倍以上
  • 适合网络运维人员和算法开发者,用于发现策略缺陷并优化性能

深度强化学习(DRL)有望实现未来移动网络的自适应控制,但传统智能体仍为被动响应:仅基于过去和当前测量值行动,无法利用带宽等外部指标的短期预测。引入预测可克服这种时间短视,但因预测增强型智能体如同黑箱,运营商难以判断决策是否受预测驱动或仅是复杂性增加,导致实际应用受限。本文提出SIA,首个能实时揭示预测增强型DRL智能体运作机制的解释器。SIA融合符号人工智能抽象与逐指标知识图谱生成解释,并引入新的影响得分指标。SIA实现亚毫秒级推理速度,较现有XAI方法快200倍以上。我们在三个不同网络场景中评估SIA,发现隐藏问题,包括预测集成的时间错位及奖励设计偏差引发反效果策略。基于这些洞察,重新设计的智能体在视频流媒体中平均码率提升9%,SIA的在线动作优化模块在不重新训练的情况下使无线接入网切片奖励提升25%。SIA使前瞻式DRL具备可解释性与可调性,降低了下一代移动网络主动控制的落地门槛。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL) promises adaptive control for future mobile networks but conventional agents remain reactive: they act on past and current measurements and cannot leverage short-term forecasts of exogenous KPIs such as bandwidth. Augmenting agents with predictions can overcome this temporal myopia, yet uptake in networking is scarce because forecast-aware agents act as closed-boxes; operators cannot tell whether predictions guide decisions or justify the added complexity. We propose SIA, the first interpreter that exposes in real time how forecast-augmented DRL agents operate. SIA fuses Symbolic AI abstractions with per-KPI Knowledge Graphs to produce explanations, and includes a new Influence Score metric. SIA achieves sub-millisecond speed, over 200x faster than existing XAI methods. We evaluate SIA on three diverse networking use cases, uncovering hidden issues, including temporal misalignment in forecast integration and reward-design biases that trigger counter-productive policies. These insights enable targeted fixes: a redesigned agent achieves a 9% higher average bitrate in video streaming, and SIA's online Action-Refinement module improves RAN-slicing reward by 25% without retraining. By making anticipatory DRL transparent and tunable, SIA lowers the barrier to proactive control in next-generation mobile networks.

强化学习可解释性网络控制预测驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。