arXiv:2511.14334cs.AIcs.LG2025-11被引 1

LLM生成约束规划模型易受文字表述影响,真实能力存疑。

When Words Change the Model: Sensitivity of LLMs for Constraint Programming Modelling

  • 通过改写经典问题描述,测试LLM对语言变化的敏感性。
  • 同一问题不同表述下,模型生成准确率显著下降。
  • 适合关注LLM可靠性与提示工程的研究者阅读。

约束规划领域长期目标是用自然语言描述问题并自动生成可执行、高效的模型。大型语言模型似乎使这一愿景更接近现实,在经典基准上展现出惊人表现。然而,这些成功可能源于数据污染而非真正推理:许多标准约束规划问题可能已存在于模型训练数据中。为验证此假设,我们系统地重述并扰动了若干知名CSPLib问题,在保持结构不变的同时改变其上下文并引入误导性元素。随后,对比了三种代表性LLM在原始与修改后描述下的模型生成结果。定性分析表明,尽管LLM能生成语法正确且语义合理的模型,但在上下文和语言变化下性能急剧下降,暴露出对表述方式的高度敏感性,揭示其理解浅层化。

原文摘要 · Abstract (English)

One of the long-standing goals in optimisation and constraint programming is to describe a problem in natural language and automatically obtain an executable, efficient model. Large language models appear to bring this vision closer, showing impressive results in automatically generating models for classical benchmarks. However, much of this apparent success may derive from data contamination rather than genuine reasoning: many standard CP problems are likely included in the training data of these models. To examine this hypothesis, we systematically rephrased and perturbed a set of well-known CSPLib problems to preserve their structure while modifying their context and introducing misleading elements. We then compared the models produced by three representative LLMs across original and modified descriptions. Our qualitative analysis shows that while LLMs can produce syntactically valid and semantically plausible models, their performance drops sharply under contextual and linguistic variation, revealing shallow understanding and sensitivity to wording.

大模型约束求解提示敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。