arXiv:2502.16556cs.CLcs.AI2025-02

测试大模型在无训练下解决管理量化问题的表现

Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving

  • 用5个主流模型在20种场景生成900条回答,测试零样本表现
  • 仅28.8%答案完全正确,复杂约束和无关参数降低准确率
  • 多步骤任务表现超预期,但重复提问无改进,模型间差异明显

本研究考察大型语言模型(LLMs)在零样本设置下解决定量管理决策问题的表现。基于五个领先模型在20种多样化管理场景中生成的900条响应,分析显示,文本呈现格式(直接、叙事或表格)及文本长度对准确性无显著影响。然而,场景复杂性——特别是约束条件和无关参数的存在——显著影响性能,常导致准确率下降。令人意外的是,模型在需要多步求解的任务中表现优于预期。仅28.8%的回复完全正确,凸显精度局限。此外,多次迭代间无显著“学习效应”,性能保持稳定。不同模型间表现差异显著,部分模型在二分类准确率上更优。总体而言,这些发现揭示了在复杂定量决策中使用LLMs的潜力与风险,为管理者和研究者提供了部署策略参考。

原文摘要 · Abstract (English)

This study examines how Large Language Models (LLMs) perform when tackling quantitative management decision problems in a zero-shot setting. Drawing on 900 responses generated by five leading models across 20 diverse managerial scenarios, our analysis explores whether these base models can deliver accurate numerical decisions under varying presentation formats, scenario complexities, and repeated attempts. Contrary to prior findings, we observed no significant effects of text presentation format (direct, narrative, or tabular) or text length on accuracy. However, scenario complexity -- particularly in terms of constraints and irrelevant parameters -- strongly influenced performance, often degrading accuracy. Surprisingly, the models handled tasks requiring multiple solution steps more effectively than expected. Notably, only 28.8\% of responses were exactly correct, highlighting limitations in precision. We further found no significant ``learning effect'' across iterations: performance remained stable across repeated queries. Nonetheless, significant variations emerged among the five tested LLMs, with some showing superior binary accuracy. Overall, these findings underscore both the promise and the pitfalls of harnessing LLMs for complex quantitative decision-making, informing managers and researchers about optimal deployment strategies.

大模型量化决策管理科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。