用大模型驱动多机器人协同,实现感知-通信-计算的智能资源调度。
Advancing Multi-Robot Networks via MLLM-Driven Sensing, Communication, and Computation: A Comprehensive Survey

- 基于大模型理解任务意图,动态分配感知、通信与计算资源。
- 四类实测系统验证:数字孪生仓储导航、主动移动控制、跟随机器人、边缘垃圾识别。
- 强调端到端性能指标,突出云端协同比纯本地处理更高效可靠。
设想由多模态大语言模型(MLLM)驱动的先进人形机器人,在仓储物流、制造及安全救援等领域协同作业。尽管单个机器人具备局部自主性,但真实任务需多个智能体共享海量传感器数据并协同工作。通信至关重要,但全量数据传输易造成网络过载,尤其当系统级编排器或云端MLLM融合多模态输入进行路径规划或异常检测时。这些任务常由高层自然语言指令触发,该意图可作为资源优化的过滤器:通过MLLM理解目标,系统可选择性激活相关传感模态,动态分配带宽,并确定计算部署位置。因此,R2X本质上是意图到资源的编排问题,需联合优化感知、通信与计算以在资源受限下最大化任务成功率。本综述探讨集成设计如何推动MLLM引导下的多机器人协同,回顾先进传感模态、通信策略与计算方法,强调推理在设备端与边缘/云端服务器间的分工。我们展示四个端到端案例:(i) 带预测链路上下文的数字孪生仓库导航,(ii) 基于移动性的主动移动计算控制,(iii) 具有语义感知切换的FollowMe机器人,(iv) 基于边缘辅助MLLM定位的真实硬件开放词汇垃圾分拣。通过载荷、延迟与成功率等系统级指标,证明R2X编排优于纯本地基线方案。
原文摘要 · Abstract (English)
Imagine advanced humanoid robots, powered by multimodal large language models (MLLMs), coordinating missions across industries like warehouse logistics, manufacturing, and safety rescue. While individual robots show local autonomy, realistic tasks demand coordination among multiple agents sharing vast streams of sensor data. Communication is indispensable, yet transmitting comprehensive data can overwhelm networks, especially when a system-level orchestrator or cloud-based MLLM fuses multimodal inputs for route planning or anomaly detection. These tasks are often initiated by high-level natural language instructions. This intent serves as a filter for resource optimization: by understanding the goal via MLLMs, the system can selectively activate relevant sensing modalities, dynamically allocate bandwidth, and determine computation placement. Thus, R2X is fundamentally an intent-to-resource orchestration problem where sensing, communication, and computation are jointly optimized to maximize task-level success under resource constraints. This survey examines how integrated design paves the way for multi-robot coordination under MLLM guidance. We review state-of-the-art sensing modalities, communication strategies, and computing approaches, highlighting how reasoning is split between on-device models and powerful edge/cloud servers. We present four end-to-end demonstrations (sense -> communicate -> compute -> act): (i) digital-twin warehouse navigation with predictive link context, (ii) mobility-driven proactive MCS control, (iii) a FollowMe robot with a semantic-sensing switch, and (iv) real-hardware open-vocabulary trash sorting via edge-assisted MLLM grounding. We emphasize system-level metrics -- payload, latency, and success -- to show why R2X orchestration outperforms purely on-device baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。