测试大模型用数学规划做决策的能力,发现进步但仍有局限。
Teaching LLMs to Think Mathematically: A Critical Study of Decision-Making via Optimization
- 用三种提示策略让大模型自动构建优化模型
- 在计算机网络问题上,最优性差距和准确率仍有提升空间
- 适合研究数学推理与人机协同决策的学者参考
本文研究大语言模型(LLMs)在使用数学规划进行决策建模与求解方面的能力。通过系统综述与元分析,评估了当前文献中LLMs在不同领域理解、构建与求解优化问题的表现,重点关注学习方法、数据集设计、评估指标与提示策略。基于新构建的数据集,针对计算机网络问题开展实验,采用‘扮演专家’、链式思考与自一致性三种提示策略,从最优性差距、分词级F1分数与编译准确率三个维度评估输出效果。结果表明,大模型在自然语言解析与符号表达方面已有显著进展,但在准确性、可扩展性与可解释性方面仍存在明显短板。这些实证差距推动了未来研究方向,包括结构化数据集、领域特定微调、混合神经符号方法、模块化多智能体架构以及基于链式RAG的动态检索。本文提出了一条推进大模型在数学规划领域能力发展的系统性路线图。
原文摘要 · Abstract (English)
This paper investigates the capabilities of large language models (LLMs) in formulating and solving decision-making problems using mathematical programming. We first conduct a systematic review and meta-analysis of recent literature to assess how well LLMs understand, structure, and solve optimization problems across domains. The analysis is guided by critical review questions focusing on learning approaches, dataset designs, evaluation metrics, and prompting strategies. Our systematic evidence is complemented by targeted experiments designed to evaluate the performance of state-of-the-art LLMs in automatically generating optimization models for problems in computer networks. Using a newly constructed dataset, we apply three prompting strategies: Act-as-expert, chain-of-thought, and self-consistency, and evaluate the obtained outputs based on optimality gap, token-level F1 score, and compilation accuracy. Results show promising progress in LLMs' ability to parse natural language and represent symbolic formulations, but also reveal key limitations in accuracy, scalability, and interpretability. These empirical gaps motivate several future research directions, including structured datasets, domain-specific fine-tuning, hybrid neuro-symbolic approaches, modular multi-agent architectures, and dynamic retrieval via chain-of-RAGs. This paper contributes a structured roadmap for advancing LLM capabilities in mathematical programming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。