测试大模型在不同表述下是否忠实遵守调度约束
SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

- 构建1132个自然语言调度题,覆盖多种场景和难度
- 13个主流大模型在不同表述下约束满足率下降明显
- 约束顺序变化对模型表现影响最显著,超出随机波动
本文提出SCHEDBench,一个基于自然语言的组合调度约束忠实度评估基准。该基准基于标准调度实例及求解器生成的可行性与最优性结果,评估大语言模型(LLMs)在不同自然语言表达形式下生成的调度方案是否保持一致的约束满足行为。SCHEDBench涵盖1,132个实例,包括作业车间调度(JSP)、单/多模式资源受限项目调度(RCPSP)、护士排班与课程安排问题,覆盖不同难度。通过领域特定模板、主题实体、词汇句法重述及约束层面的表面形式变异,将实例转化为自然语言问题,并验证参考解的可行性和目标最优性。在13个前沿及开源大模型上测试发现,模型对语义等价但表述不同的问题缺乏一致性;表面形式变化导致可行性下降,并引发超过噪声水平的硬约束违规变化。在各独立变量中,约束重排产生的敏感度最为显著。
原文摘要 · Abstract (English)
This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility and optimality, SCHEDBench assesses whether large language models (LLMs) generate schedules with the same constraint-feasible behavior across varied natural-language (NL) surface forms. SCHEDBench spans 1,132 instances across job-shop scheduling problems (JSP), single and multi-mode resource-constrained project scheduling problems (RCPSP), nurse rostering/scheduling, and curriculum timetabling problems of varying difficulty. Instances are templated into natural language problems using domain-specific templates, themed entities, lexical-syntactic template rephrasing, and constraint-level surface-form variation, with reference solutions verified for feasibility and objective optimality. Across thirteen frontier and open-weight LLMs, we find that models are not reliably invariant to semantically equivalent renderings of the same scheduling problem. Surface-form variation reduces feasibility and induces above-noise shifts in per-instance hard-constraint violations on matched instances. Among the tested isolated axes, constraint reordering yields the clearest above-noise sensitivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。