arXiv:2603.14729cs.LGcs.DC2026-03

用去中心化强化学习实现跨孤岛物联网调度,兼顾效率与安全。

DeFRiS: Silo-Cooperative IoT Applications Scheduling via Decentralized Federated Reinforcement Learning

  • 通过无动作空间依赖的策略,实现异构节点间知识无缝迁移。
  • 在20个异构节点上降低6.4%响应时间,尾延迟风险减少10.4%。
  • 适合大规模、存在恶意攻击的分布式物联网系统部署。

下一代物联网应用日益跨越多个自治实体,亟需在保护数据隐私的前提下实现跨孤岛协作调度以利用多样化计算资源。然而,基础设施异构性、非独立同分布的工作负载变化以及对抗环境固有风险,使高效协作面临挑战。现有方法多依赖中心化协调或独立学习,难以解决异构孤岛间状态-动作空间不兼容问题,且对恶意攻击缺乏鲁棒性。本文提出DeFRiS,一种用于鲁棒可扩展孤岛协同物联网应用调度的去中心化联邦强化学习框架。其融合三项创新:(i) 无动作空间依赖的策略,基于候选资源评分实现异构孤岛间无缝知识转移;(ii) 孤岛优化的本地学习机制,结合广义优势估计(GAE)与截断策略更新,缓解稀疏延迟奖励难题;(iii) 双轨非独立同分布鲁棒去中心化聚合协议,利用梯度指纹实现相似性感知的知识迁移与异常检测,并通过梯度追踪保持优化动量。在包含20个异构孤岛的真实物联网工作负载测试平台上,实验表明DeFRiS显著优于现有最优基线,平均响应时间降低6.4%,能耗减少7.2%,尾延迟风险(CVaR$_{0.95}$)下降10.4%,几乎零截止时间违规。此外,系统扩展时性能保留超过3倍,对抗环境中稳定性提升超8倍。

原文摘要 · Abstract (English)

Next-generation IoT applications increasingly span across autonomous administrative entities, necessitating silo-cooperative scheduling to leverage diverse computational resources while preserving data privacy. However, realizing efficient cooperation faces significant challenges arising from infrastructure heterogeneity, Non-IID workload shifts, and the inherent risks of adversarial environments. Existing approaches, relying predominantly on centralized coordination or independent learning, fail to address the incompatibility of state-action spaces across heterogeneous silos and lack robustness against malicious attacks. This paper proposes DeFRiS, a Decentralized Federated Reinforcement Learning framework for robust and scalable Silo-cooperative IoT application scheduling. DeFRiS integrates three synergistic innovations: (i) an action-space-agnostic policy utilizing candidate resource scoring to enable seamless knowledge transfer across heterogeneous silos; (ii) a silo-optimized local learning mechanism combining Generalized Advantage Estimation (GAE) with clipped policy updates to resolve sparse delayed reward challenges; and (iii) a Dual-Track Non-IID robust decentralized aggregation protocol leveraging gradient fingerprints for similarity-aware knowledge transfer and anomaly detection, and gradient tracking for optimization momentum. Extensive experiments on a distributed testbed with 20 heterogeneous silos and realistic IoT workloads demonstrate that DeFRiS significantly outperforms state-of-the-art baselines, reducing average response time by 6.4% and energy consumption by 7.2%, while lowering tail latency risk (CVaR$_{0.95}$) by 10.4% and achieving near-zero deadline violations. Furthermore, DeFRiS achieves over 3 times better performance retention as the system scales and over 8 times better stability in adversarial environments compared to the best-performing baseline.

物联网调度联邦学习强化学习去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。