用外部求解器增强大模型,高效解决运筹学问题
OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
- 用半自动数据生成+外部求解器增强模型
- 在4个基准上最高达80.1%执行准确率
- 零样本泛化能力提升21个百分点
大语言模型虽具强数学推理能力,但依赖闭源API处理运筹学任务存在隐私风险,而从头训练开源模型成本过高。我们提出OR-Toolformer,基于Llama-3.1-8B-Instruct,通过半自动数据合成管道生成多样化的运筹学问题-答案对,并引入外部求解器以生成API调用。在四个标准基准中的三个上,其执行准确率最高达80.1%,优于同规模基线超过4.3%。在两个未见过的运筹学问题类型上进行零样本评估时,平均准确率达54%,较最强基线提升21个百分点。结果验证了工具增强微调大模型在运筹学问题建模与求解中兼具准确性与泛化性的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate strong mathematical reasoning, but reliance on closed-source APIs for OR tasks raises privacy concerns, and training open-source models from scratch incurs high compute costs. We introduce OR-Toolformer, which fine-tunes Llama-3.1-8B-Instruct with a semi-automatic data synthesis pipeline that generates diverse OR problem-answer pairs and augments the model with external solvers to produce API calls. On three of four standard benchmarks, OR-Toolformer achieves up to 80.1% execution accuracy, exceeding size-matched baselines by over 4.3%. In zero-shot evaluation on two unseen OR problem types, it attains 54% average accuracy, a 21 percentage-point improvement over the strongest baseline. These findings validate the efficacy of tool-augmented fine-tuning LLMs for accurate and generalizable OR problem modeling and solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。