arXiv:2603.10436cs.ROcs.DC2026-03中稿 · 27th IEEE Internat…

多机器人协作推理框架,实时优化算力分配,省电51%且提速2.5倍。

COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints

  • 混合强化学习动态分配模型计算任务,分担各机器人负载。
  • 实测电池消耗降15.4%,GPU利用率提升51.67%,满足2.55倍帧率要求。
  • 适用于灾后救援等无网络支持的多机协同场景,适合边缘智能研究者。

大型深度神经网络(如基于Transformer和多模态架构)在资源受限的边缘平台(如野外机器人)上部署困难。在灾后救援等关键任务中,机器人需在带宽、延迟和电池寿命严格受限下协作,且缺乏基础设施支持。为此,我们提出COHORT——基于ROS的多机器人协同DNN推理与任务执行框架。该框架采用混合离线-在线强化学习策略,动态调度并分配各机器人间的DNN模块执行任务。主要贡献包括:(a) 基于拍卖式任务分配数据,利用优势加权回归(AWR)训练离线强化学习策略;(b) 以离线策略为初始,通过多智能体PPO(MAPPO)在线实时微调;(c) 在视觉语言模型(如CLIP、SAM)推理任务上全面评估其可扩展性与鲁棒性。实验对比遗传算法及多种强化学习基线,结果表明:相比基准,COHORT降低15.4%电池消耗,提升51.67% GPU利用率,并在2.55倍时间内满足帧率与截止时间约束。

原文摘要 · Abstract (English)

Large deep neural networks (DNNs), especially transformer-based and multimodal architectures, are computationally demanding and challenging to deploy on resource-constrained edge platforms like field robots. These challenges intensify in mission-critical scenarios (e.g., disaster response), where robots must collaborate under tight constraints on bandwidth, latency, and battery life, often without infrastructure or server support. To address these limitations, we present COHORT, a collaborative DNN inference and task-execution framework for multi-robot systems built on the Robotic Operating System (ROS). COHORT employs a hybrid offline-online reinforcement learning (RL) strategy to dynamically schedule and distribute DNN module execution across robots. Our key contributions are threefold: (a) Offline RL policy learning combined with Advantage-Weighted Regression (AWR), trained on auction-based task allocation data from heterogeneous DNN workloads across distributed robots, (b) Online policy adaptation via Multi-Agent PPO (MAPPO), initialized from the offline policy and fine-tuned in real time, and (c) comprehensive evaluation of COHORT on vision-language model (VLM) inference tasks such as CLIP and SAM, analyzing scalability with increasing robot/workload and robustness under . We benchmark COHORT against genetic algorithms and multiple RL baselines. Experimental results demonstrate that COHORT reduces battery consumption by 15.4% and increases GPU utilization by 51.67%, while satisfying frame-rate and deadline constraints 2.55 times of the time.

多机器人边缘计算强化学习协同推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。