用局部错误驱动方法提升大模型优化建模能力,效果显著优于现有技术。
Automated Optimization Modeling via a Localizable Error-Driven Perspective
- 基于错误传播的局部性,构建聚焦高密度训练数据
- 在6个基准上超越当前最优方法,解决难例泛化问题
- 适合需要自动化建模的工业场景与复杂决策任务
利用大语言模型(LLMs)实现自动化优化建模已成为辅助复杂人类决策的有前景方向。尽管后训练已成为提升该领域能力的关键技术,但其效果受限于高质量训练数据的稀缺与低效利用。通过对后训练中多种问题-回答对的误差模式深入分析,我们识别出现有方法的两大根本缺陷:(L1) 错误相关问题稀疏,(L2) 难题对应奖励稀疏。实验表明,这些限制会导致领域特定后训练中性能次优。为此,我们提出一种新型误差驱动学习框架——基于可定位误差驱动视角的自动化优化建模(MIND),从数据合成到后训练全流程定制。核心洞察是:优化建模中的错误具有独特可定位性,即错误常局限于特定语义片段,不会在整个解中传播。因此,与数学证明等整体推理任务不同,MIND通过构建聚焦、高密度训练语料,并提出动态监督微调策略优化(DFPO),以局部精炼方式攻克难题。六项基准测试结果表明,MIND始终优于所有现有先进方法。
原文摘要 · Abstract (English)
Automated optimization modeling via Large Language Models (LLMs) has emerged as a promising approach to assist complex human decision-making. While post-training has become a pivotal technique to enhance LLMs' capabilities in this domain, its effectiveness is severely constrained by the scarcity and underutilization of high-quality training data. However, through a detailed profiling of error patterns across various problem-response pairs drawn from post-training, we identify two fundamental limitations of existing automated optimization modeling approaches: (L1) the sparsity of error-specific problems and (L2) the sparse rewards associated with difficult problems. We demonstrate that these limitations can result in suboptimal performance in domain-specific post-training for LLMs. To tackle the above two limitations, we propose a novel error-driven learning framework -- namely, auto\textbf{m}ated opt\textbf{i}mization modeli\textbf{n}g via a localizable error-\textbf{d}riven perspective (MIND) -- that customizes the whole model training framework from data synthesis to post-training. MIND is based on our key observation of the unique localizable patterns in error propagation of optimization modelings, that is, modeling errors may remain localized to specific semantic segments and do not propagate throughout the entire solution. Thus, in contrast to holistic reasoning tasks such as mathematical proofs, MIND leverages the construction of a focused, high-density training corpus and proposes \textbf{D}ynamic Supervised \textbf{F}ine-Tuning \textbf{P}olicy \textbf{O}ptimization (DFPO) to tackle difficult problems through localized refinement. Experiments on six benchmarks demonstrate that MIND consistently outperforms all the state-of-the-art automated optimization modeling approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。