arXiv:2512.03955cs.AIcs.ET2025-12被引 2

为大模型智能体规划控制设计标准化测试基准

Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol

  • 构建可执行的积木世界仿真环境,支持五类复杂度任务
  • 通过统一接口协议连接不同智能体架构,实现无修改评估
  • 提供量化指标,适合研究大模型自主决策与执行的学者

工业自动化日益需要能适应变化任务和环境的灵活控制策略。基于大语言模型(LLMs)的智能体在自适应规划与执行方面具有潜力,但缺乏标准化的评估基准。本文提出一个基于积木世界问题的可执行仿真环境,包含五个复杂度等级。通过集成模型上下文协议(Model Context Protocol, MCP)作为标准化工具接口,使多种智能体架构无需修改即可接入并评估。单一智能体实现验证了该基准的可用性,建立了用于比较基于LLM的规划与执行方法的定量指标。

原文摘要 · Abstract (English)

Industrial automation increasingly requires flexible control strategies that can adapt to changing tasks and environments. Agents based on Large Language Models (LLMs) offer potential for such adaptive planning and execution but lack standardized benchmarks for systematic comparison. We introduce a benchmark with an executable simulation environment representing the Blocksworld problem providing five complexity categories. By integrating the Model Context Protocol (MCP) as a standardized tool interface, diverse agent architectures can be connected to and evaluated against the benchmark without implementation-specific modifications. A single-agent implementation demonstrates the benchmark's applicability, establishing quantitative metrics for comparison of LLM-based planning and execution approaches.

大模型智能体规划控制基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。