arXiv:2503.10642cs.CLcs.AI2025-03被引 13

构建跨领域自然语言优化问题数据集,助力大模型建模求解。

Text2Zinc: A Cross-Domain Dataset for Modeling Optimization and Satisfaction Problems in MiniZinc

  • 用MiniZinc统一建模自然语言描述的优化与满足问题。
  • 实验表明大模型直接建模仍不成熟,需改进方法。
  • 适合研究大模型辅助约束求解的学者使用。

近年来,利用大语言模型(LLMs)作为组合优化和约束编程任务的协作者日益受到关注。本文提出Text2Zinc,一个跨领域的数据集,用于捕捉以自然语言文本描述的优化与满足问题。与以往工作不同,本研究在统一数据集中整合了优化与满足问题,并采用与求解器无关的建模语言进行表达。我们借助MiniZinc的求解器与范式无关建模能力实现问题形式化。基于Text2Zinc,我们开展全面基线实验,比较多种方法在执行效率与解准确率上的表现,包括现成提示策略、思维链推理及组合式方法。此外,还探索了中间表示(如知识图谱)的有效性。结果表明,目前大模型尚不能作为“一键式”工具从文本中建模组合问题。我们期望Text2Zinc能为研究人员和实践者提供重要资源,推动该领域进一步发展。

原文摘要 · Abstract (English)

There is growing interest in utilizing large language models (LLMs) as co-pilots for combinatorial optimization and constraint programming tasks across various problems. This paper aims to advance this line of research by introducing Text2Zinc}, a cross-domain dataset for capturing optimization and satisfaction problems specified in natural language text. Our work is distinguished from previous attempts by integrating both satisfaction and optimization problems within a unified dataset using a solver-agnostic modeling language. To achieve this, we leverage MiniZinc's solver-and-paradigm-agnostic modeling capabilities to formulate these problems. Using the Text2Zinc dataset, we conduct comprehensive baseline experiments to compare execution and solution accuracy across several methods, including off-the-shelf prompting strategies, chain-of-thought reasoning, and a compositional approach. Additionally, we explore the effectiveness of intermediary representations, specifically knowledge graphs. Our findings indicate that LLMs are not yet a push-button technology to model combinatorial problems from text. We hope that Text2Zinc serves as a valuable resource for researchers and practitioners to advance the field further.

大模型约束求解自然语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。