构建真实世界数学建模评估框架,让大模型解决复杂开放问题。
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
- 设计多智能体系统,协调工具使用与迭代优化求解过程。
- 在真实竞赛题上超越基线模型,结果接近人类专家水平。
- 支持多种合理解法,适合研究跨学科建模与智能决策的人群。
大型语言模型在解决数学问题上取得显著进展,但现有基准难以反映真实世界的复杂性,后者需要开放式的跨学科推理和计算工具整合。为此,我们提出ModelingBench,一个基于真实数学建模竞赛的新型基准,涵盖从城市交通优化到生态系统资源规划等多领域开放性问题。这些任务要求将自然语言转化为形式化数学表达、调用适当工具并生成结构化、可辩护的报告。ModelingBench允许多种有效解法,体现实际建模中的模糊性与创造性。同时,我们提出ModelingAgent——一种多智能体框架,通过协调工具使用、支持结构化工作流并实现迭代自我修正,生成有依据且具创造性的解决方案。为评估输出,我们进一步设计ModelingJudge,一个专家参与的闭环系统,利用领域专精的LLM从多角度评判解决方案。实证结果显示,ModelingAgent显著优于强基线,常产出与人类专家难以区分的结果。本工作共同提供了一个评估与推进开放式跨学科建模能力的完整框架。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity of real-world problems, which demand open-ended, interdisciplinary reasoning and integration of computational tools. To address this gap, we introduce ModelingBench, a novel benchmark featuring real-world-inspired, open-ended problems from math modeling competitions across diverse domains, ranging from urban traffic optimization to ecosystem resource planning. These tasks require translating natural language into formal mathematical formulations, applying appropriate tools, and producing structured, defensible reports. ModelingBench also supports multiple valid solutions, capturing the ambiguity and creativity of practical modeling. We also present ModelingAgent, a multi-agent framework that coordinates tool use, supports structured workflows, and enables iterative self-refinement to generate well-grounded, creative solutions. To evaluate outputs, we further propose ModelingJudge, an expert-in-the-loop system leveraging LLMs as domain-specialized judges assessing solutions from multiple expert perspectives. Empirical results show that ModelingAgent substantially outperforms strong baselines and often produces solutions indistinguishable from those of human experts. Together, our work provides a comprehensive framework for evaluating and advancing real-world problem-solving in open-ended, interdisciplinary modeling challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。