arXiv:2510.05486cs.CL2025-10ACL被引 3

给大模型规划任务添加精细约束,发现性能普遍下降一半。

Language Model as Planner and Formalizer under Constraints

  • 在经典规划基准上加入四类细粒度自然语言约束
  • 4个主流推理模型在加约束后性能平均下降50%
  • 揭示现有大模型规划能力脆弱性,适合安全敏感场景研究

大模型在规划任务中既可作为端到端生成动作序列的规划器,也可作为将规划领域和问题形式化为确定性计划语言的转换器。然而,当前方法依赖仅含通用简单环境描述的标准基准,可能导致对大模型规划能力的高估,并引发下游应用中的安全风险。本文通过人工标注、细粒度且丰富的自然语言约束(涵盖四类形式化类别)增强常用规划基准,覆盖4个前沿推理大模型、4种形式化语言及4个数据集。结果表明,引入单句约束后性能普遍下降约50%,凸显当前大模型在约束条件下的脆弱性,也为未来研究提供方向。

原文摘要 · Abstract (English)

LLMs have been widely used in planning, either as planners to generate action sequences end-to-end, or as formalizers to represent the planning domain and problem in a formal language that can derive plans deterministically. However, both lines of work rely on standard benchmarks that include only generic and simplistic environmental specifications, leading to potential overestimation of the planning ability of LLMs and safety concerns in downstream tasks. We bridge this gap by augmenting widely used planning benchmarks with manually annotated, fine-grained, and rich natural language constraints spanning four formally defined categories. Over 4 state-of-the-art reasoning LLMs, 4 formal languages, and 4 datasets, we show that the introduction of one-sentence constraints consistently halves performance, indicating current LLMs' lack of robustness and an avenue for future research.

大模型规划约束推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。