提出工业级多智能体控制新基准,验证专业化与集中化权衡
Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control
- 融合排序与压榨任务构建序列化回收场景
- 动作掩码使双架构性能提升,专业化优势减弱
- 适合研究工业自动化中的多智能体强化学习
多阶段工业过程的自主控制需兼顾局部专业化与全局协调。强化学习虽具潜力,但受限于奖励设计、模块化与动作空间管理等问题,工业应用仍有限。现有学术基准与真实工业问题差异显著,制约了可迁移性。本文提出一个增强型工业启发式基准环境,整合SortingEnv与ContainerGym两个基准的任务,构建包含分拣与压榨操作的序列化回收场景。评估两种控制策略:专用智能体的模块化架构与统一管控全系统的单体智能体,并分析动作掩码的影响。实验表明,无动作掩码时,智能体难以学习有效策略,模块化架构表现更优;引入动作掩码后,两类架构性能均显著提升,性能差距大幅缩小。结果凸显动作空间约束的关键作用,表明当动作复杂度降低时,专业化优势减弱。该基准为探索工业自动化中实用且鲁棒的多智能体强化学习方案提供了重要测试平台,同时推动了集中化与专业化之争的讨论。
原文摘要 · Abstract (English)
Autonomous control of multi-stage industrial processes requires both local specialization and global coordination. Reinforcement learning (RL) offers a promising approach, but its industrial adoption remains limited due to challenges such as reward design, modularity, and action space management. Many academic benchmarks differ markedly from industrial control problems, limiting their transferability to real-world applications. This study introduces an enhanced industry-inspired benchmark environment that combines tasks from two existing benchmarks, SortingEnv and ContainerGym, into a sequential recycling scenario with sorting and pressing operations. We evaluate two control strategies: a modular architecture with specialized agents and a monolithic agent governing the full system, while also analyzing the impact of action masking. Our experiments show that without action masking, agents struggle to learn effective policies, with the modular architecture performing better. When action masking is applied, both architectures improve substantially, and the performance gap narrows considerably. These results highlight the decisive role of action space constraints and suggest that the advantages of specialization diminish as action complexity is reduced. The proposed benchmark thus provides a valuable testbed for exploring practical and robust multi-agent RL solutions in industrial automation, while contributing to the ongoing debate on centralization versus specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。