用提示修正提升大模型优化建模能力,效果接近千亿参数模型
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
- 通过专家提示纠正推理错误,仅修改2.6%文本实现高效微调
- 40亿参数模型在5个基准上达68.9%准确率,媲美6710亿参数模型
- 适合需要高精度建模的科研与工程人员使用
大型推理模型(LRMs)在多步复杂推理中表现出强大能力,为自动化优化建模带来新机遇。然而,现有领域适应方法针对早期指令微调模型设计,难以发挥现代LRMs的先进推理模式——我们发现,直接在传统非反思性数据集上微调效果有限。为此,我们提出CALM(Corrective Adaptation with Lightweight Modification)框架,通过专家干预识别推理缺陷并提供简洁修正提示,使LRM在原始推理模式下逐步优化。干预仅修改生成文本的2.6%以下,却生成高质量数据用于监督微调。随后模型通过强化学习进一步优化。基于CALM,我们构建了40亿参数的STORM(Smart Thinking Optimization Reasoning Model),在五个主流优化建模基准上达到68.9%的平均准确率,与6710亿参数模型性能相当。结果表明,动态提示驱动的数据合成既能保留又能放大现代LRMs的原生推理能力,为实现挑战性优化建模任务的专家级表现提供了更有效、可扩展的路径。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have demonstrated strong capabilities in complex multi-step reasoning, opening new opportunities for automating optimization modeling. However, existing domain adaptation methods, originally designed for earlier instruction-tuned models, often fail to exploit the advanced reasoning patterns of modern LRMs -- In particular, we show that direct fine-tuning on traditional \textit{non-reflective} datasets leads to limited gains. To fully leverage LRMs' inherent reasoning abilities, we propose \textbf{CALM} (\textit{Corrective Adaptation with Lightweight Modification}), a framework that progressively refines LRMs within their native reasoning modes for optimization modeling tasks. In CALM, an expert intervener identifies reasoning flaws and provides concise corrective hints, which the LRM incorporates to produce improved reasoning trajectories. These interventions modify fewer than 2.6\% of generated tokens, but generate high-quality data for soft adaptation through supervised fine-tuning. The adapted model is then further improved through reinforcement learning. Building on CALM, we develop \textbf{STORM} (\textit{Smart Thinking Optimization Reasoning Model}), a 4B-parameter LRM that achieves a new state-of-the-art average accuracy of 68.9\% across five popular optimization modeling benchmarks, matching the performance of a 671B LRM. These results demonstrate that dynamic, hint-based data synthesis both preserves and amplifies the native reasoning patterns of modern LRMs, offering a more effective and scalable path towards expert-level performance on challenging optimization modeling tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。