arXiv:2608.07148cs.AI2026-08

用多智能体强化学习框架,设计大模型在智能制造中的协同控制集成方案

A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing

  • 以集中训练、分散执行的马尔可夫决策过程为基底,构建四类大模型接入点
  • 实证表明大模型更适合语义理解与慢速规划,传统MARL更适于高频结构化协同
  • 提出三层参考架构,支持可验证执行,适合工业级智能系统开发

现代制造对自适应控制提出六大耦合需求:局部决策产生全局影响、部分可观测、非平稳性、快速响应与长周期效应并存、延迟且分散的结果反馈,以及难以显式建模的动态特性。协作式多智能体强化学习(MARL)以集中训练、分散执行的Dec-POMDP形式,天然契合这些挑战。本文聚焦于MARL框架,探讨大语言模型(LLM)应在何处增强、接口、训练或替代其核心协调机制。通过四类LLM接入点(策略、奖励设计、智能体间通信、层级规划)构建文献分类体系;基于能力成熟度分析,区分原生机制、性能表现、形式保证与工程成熟度,并评估部署可行性。最终贡献是一个三层证据驱动的参考架构,支持语义推理、自适应协同控制和独立可验证执行。提出的LLM增强型Dec-POMDP仅用于描述不同接入方式,不引入新算法。现有证据显示,传统MARL更适合任务特定训练后的高频、结构化、去中心化协调;而大模型在语义解释、奖励设计、人机交互及慢速监督规划中具潜力。当前大模型尚未证明能等效替代实时、去中心化、安全关键的制造控制器,此结论基于现有证据,不否定未来可能性。

原文摘要 · Abstract (English)

Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nonstationarity, reflex speed response with long horizon effects, delayed and diffuse outcomes, and dynamics that resist explicit modeling. Cooperative multiagent reinforcement learning (MARL), posed as a Dec-POMDP under centralized training with decentralized execution, is a particularly natural formalism for these demands. This paper adopts a MARL centered scope and asks where large language models (LLMs) should augment, interface with, train, or, in the strongest competitive case, replace that coordination core. A taxonomy organizes the literature through four LLM attachment points: policy, reward design, communication between agents, and hierarchical planning. A conditional capability profile separates native mechanism, reported performance, formal guarantee, and engineering maturity, and a deployment readiness analysis identifies the evidence behind each role. These stages yield the principal contribution: a three layer MARL centered reference architecture, grounded in evidence, for semantic reasoning, adaptive cooperative control, and independently assured execution. The LLM-Augmented Dec-POMDP is a descriptive comparative notation for that architecture, recording four attachment choices without introducing a new decision process class or algorithm. Under the reviewed evidence, conventional MARL is better suited to frequent, structured, decentralized coordination after task specific training, whereas LLM components are promising for semantic interpretation, reward drafting, human interaction, and slower supervisory planning. Current LLM only manufacturing controllers do not yet establish equivalence for strict real time, decentralized, safety critical control; this conclusion is bounded by the available evidence and does not assert impossibility.

多智能体大模型智能制造强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。