用真实运筹学问题测试大模型推理能力,发现其表现有限。
Evaluating LLM Reasoning in the Operations Research Domain with ORQA
- 构建运筹学问答数据集,要求多步建模推理
- 主流大模型在真实问题上表现平庸,泛化能力不足
- 适合研究大模型在专业领域泛化能力的学者
本文提出并应用了运筹学问答(ORQA)基准,用于评估大语言模型(LLMs)在运筹学这一专业领域的泛化能力。该基准旨在检验大模型面对多样化复杂优化问题时,能否具备运筹学专家的知识与推理能力。数据集由运筹学专家构建,包含需多步推理才能建立数学模型的真实世界优化问题。我们对多种开源大模型(如LLaMA 3.1、DeepSeek、Mixtral)进行评估,结果显示其表现有限,暴露出在专业领域泛化能力上的明显差距。本工作推动了关于大模型泛化能力的讨论,为后续研究提供重要参考。数据集与评估代码已公开。
原文摘要 · Abstract (English)
In this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark designed to assess the generalization capabilities of Large Language Models (LLMs) in the specialized technical domain of Operations Research (OR). This benchmark evaluates whether LLMs can emulate the knowledge and reasoning skills of OR experts when confronted with diverse and complex optimization problems. The dataset, developed by OR experts, features real-world optimization problems that demand multistep reasoning to construct their mathematical models. Our evaluations of various open source LLMs, such as LLaMA 3.1, DeepSeek, and Mixtral, reveal their modest performance, highlighting a gap in their ability to generalize to specialized technical domains. This work contributes to the ongoing discourse on LLMs generalization capabilities, offering valuable insights for future research in this area. The dataset and evaluation code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。