arXiv:2505.23881cs.AIcs.CL2025-05被引 7

用推理模型生成搜索策略,破解多个组合设计难题

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems

  • 通过推理型大模型生成搜索策略,自动优化求解过程
  • 成功解决16个组合设计问题中的7个长期未解难题
  • 适合对数学求解、AI辅助科研感兴趣的读者

具有推理能力的大语言模型(LLMs)可通过迭代生成与优化答案,在数学和代码生成中表现优异。本文将此类模型应用于组合设计领域,该领域存在多种尚未确定存在性的开放实例。提出构造协议CPro1,基于文本定义与有效性验证器,引导大模型选择并实现策略,结合自动化超参数调优与执行反馈。使用推理型LLMs的CPro1成功解决了《组合设计手册》(2006)中16个问题里的7个长期未解实例,包括布莱斯卡·罗设计、对称加权矩阵、平衡三元设计等3类问题,这些此前在非推理模型下未能解决。此外还解决了2025年新文献中的若干开放问题,生成了新的覆盖序列、约翰逊团覆盖、删除码及均匀嵌套斯坦纳四元系。

原文摘要 · Abstract (English)

Large Language Models (LLMs) with reasoning are trained to iteratively generate and refine their answers before finalizing them, which can help with applications to mathematics and code generation. We apply code generation with reasoning LLMs to a specific task in the mathematical field of combinatorial design. This field studies diverse types of combinatorial designs, many of which have lists of open instances for which existence has not yet been determined. The Constructive Protocol CPro1 uses LLMs to generate search heuristics that have the potential to construct solutions to small open instances. Starting with a textual definition and a validity verifier for a particular type of design, CPro1 guides LLMs to select and implement strategies, while providing automated hyperparameter tuning and execution feedback. CPro1 with reasoning LLMs successfully solves long-standing open instances for 7 of 16 combinatorial design problems selected from the 2006 Handbook of Combinatorial Designs, including new solved instances for 3 of these (Bhaskar Rao Designs, Symmetric Weighing Matrices, Balanced Ternary Designs) that were unsolved by CPro1 with non-reasoning LLMs. It also solves open instances for several problems from recent (2025) literature, generating new Covering Sequences, Johnson Clique Covers, Deletion Codes, and a Uniform Nested Steiner Quadruple System.

组合设计推理模型数学求解AI科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。