评测大模型解离散优化问题能力,发现思维链未必有效
Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning
- 用多样化数据集测试大模型求解离散优化问题
- 强模型表现更好,但思维链效果不总是提升
- 适合想自动化解决优化问题的研究者参考
本研究评估了Llama-3系列与CHATGPT等大模型在离散优化问题上的表现,使用包含多种问题类型和大规模参数的自然语言数据集。数据集分为原始、扩展和增强三类:原始与增强用于评估,扩展可用于微调新模型。实验对比了强弱模型、带与不带思维链(CoT)方法在不同数据集上的表现。结果显示,强模型性能更优;但出人意料的是,思维链并非始终有效;且杂乱数据反而提升了模型在易理解问题上的表现,尽管存在较高方差,反映不稳定性。研究旨在全面评估大模型在大规模优化问题中的能力,为自动求解提供建议,并为后续研究提供基准。附录中提供了详细图表以供参考。
原文摘要 · Abstract (English)
This work investigated the capabilities of different models, including the Llama-3 series of models and CHATGPT, with different forms of expression in solving discrete optimization problems by testing natural language datasets. In contrast to formal datasets with a limited scope of parameters, our dataset included a variety of problem types in discrete optimization problems and featured a wide range of parameter magnitudes, including instances with large parameter sets, integrated with augmented data. It aimed to (1) provide an overview of LLMs' ability in large-scale problems, (2) offer suggestions to those who want to solve discrete optimization problems automatically, and (3) regard the performance as a benchmark for future research. These datasets included original, expanded and augmented datasets. Among these three datasets, the original and augmented ones aimed for evaluation while the expanded one may help finetune a new model. In the experiment, comparisons were made between strong and week models, CoT methods and No-CoT methods on various datasets. The result showed that stronger model performed better reasonably. Contrary to general agreement, it also showed that CoT technique was not always effective regarding the capability of models and disordered datasets improved performance of models on easy to-understand problems, even though they were sometimes with high variance, a manifestation of instability. Therefore, for those who seek to enhance the automatic resolution of discrete optimization problems, it is recommended to consult the results, including the line charts presented in the Appendix, as well as the conclusions drawn in this study for relevant suggestions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。