arXiv:2609.06896cs.RO2026-09

提出分布式安全学习控制框架,应对多机器人系统隐蔽执行器攻击。

Distributed Secure Learning Control for Large-scale Multirobots under Stealthy Actuator Attacks

论文配图:Distributed Secure Learning Control for Large-scale Multirobots under Stealthy Actuator Attacks
图 1 · 摘自论文原文
  • 基于博弈论的分布式强化学习架构,实时学习攻防策略。
  • 在多种攻击场景下保持控制性能,支持大规模机器人系统。
  • 适用于有隐蔽攻击风险的多机器人协同任务,如自主导航。

多机器人系统(MRS)的分布式学习控制虽具灵活性,但缺乏可证明的性能保证。将强化学习(RL)与分布式模型预测控制(DMPC)结合,可利用RL在非线性策略设计中的优势及DMPC的滚动重规划能力。然而,在恶意网络攻击(尤其是隐蔽攻击)下保障控制安全仍是关键挑战,因分布式策略依赖邻近节点间的信息交换,被攻陷的节点可通过通信网络快速影响其他节点行为。本文提出一种针对大规模多机器人系统在隐蔽执行器攻击下的分布式安全学习控制(DSLC)框架。该框架具备两个核心特性:(i) 统一方法,适用于多种协调场景;(ii) 基于微分博弈的分布式学习型预测控制策略,通过学习攻防平衡机制实现鲁棒控制。具体而言,DSLC采用分布式攻击者-行动者-评论家架构,在每个预测区间内在线学习最优防御与攻击策略。不同于依赖数值优化生成开环控制序列的方法,本方法以解析闭式形式同时生成对抗性攻击策略与对应防御策略。防御策略可直接推广至不同规模和不同执行器攻击概率的多机器人系统。通过多轮轮式机器人仿真与真实实验,在多种控制任务中验证了DSLC的有效性与可扩展性。

原文摘要 · Abstract (English)

Distributed learning control for multirobot systems (MRS) offers significant flexibility in presence of uncertainties but lacks provable performance guarantees. A promising direction involves integrating reinforcement learning (RL) into distributed model predictive control (DMPC), leveraging the strengths of RL in nonlinear policy design and the receding-horizon replanning capabilities of DMPC. However, ensuring secure control within such a learning framework under malicious cyber attacks, particularly stealthy ones, remains a critical challenge, because the distributed policies generation depends on information exchange among neighbors, where compromised agents can rapidly influence the behavior of others through the communication network. This article proposes a distributed secure learning control (DSLC) framework for large-scale MRS under malicious, stealthy actuator attacks. Our framework offers two key features: (i) a unified approach that enables secure learning control across various coordination scenarios and (ii) a game-theoretic distributed learning-based predictive control strategy that learns how to balance the attacker and defender through a differential-game based DMPC framework. Specifically, DSLC employs a distributed attacker-actor-critic architecture to learn the optimal defense and attack policies online within each prediction interval. Unlike numerical optimization-based controllers that calculate open-loop control sequences, our method simultaneously generates adversarial attack policies and corresponding defense policies in analytical closed-loop form. The defense policies could be directly generalized to MRS with varying scales and diverse actuator attack probabilities. The effectiveness and scalability of DSLC are validated through comprehensive simulations and real-world experiments in multiple wheeled robots via various control tasks.

多机器人安全控制强化学习攻防博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。