arXiv:2506.00577cs.AIcs.CL2025-06被引 1

用经济学问题微调大模型,让其学会策略性思考。

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs

  • 在2100个高质量经济题上微调7B模型,提升结构化推理能力。
  • 在多智能体博弈中表现更理性,决策更符合经济规律。
  • 适合研究模型推理、智能体对齐及经济应用的开发者。

直接训练大语言模型(LLMs)用于多智能体系统(MAS)仍具挑战,受限于复杂的奖励建模、动态交互和泛化需求。本文探索后训练技术(如监督微调SFT和可验证奖励强化学习RLVR)在多智能体场景中的泛化效果。以经济推理为测试基准,因其具备数学与博弈论基础、强结构化分析要求,且适用于市场设计、资源分配等真实场景。我们提出Recon(Rationality via Economics Reasoning),一个7B参数的开源模型,在2100个精心筛选的高质量经济推理问题上进行后训练。在经济推理基准和多智能体游戏上的全面评估显示,模型在结构化推理和经济理性方面均有显著提升。结果表明,领域对齐的后训练能有效增强推理与智能体对齐,揭示了SFT与RL在塑造模型行为中的作用。代码已开源:https://github.com/MasterZhou1/Recon。

原文摘要 · Abstract (English)

Directly training Large Language Models (LLMs) for Multi-Agent Systems (MAS) remains challenging due to intricate reward modeling, dynamic agent interactions, and demanding generalization requirements. This paper explores whether post-training techniques, specifically Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR), can effectively $\textit{generalize}$ to multi-agent scenarios. We use economic reasoning as a testbed, leveraging its strong foundations in mathematics and game theory, its demand for structured analytical reasoning, and its relevance to real-world applications such as market design, resource allocation, and policy analysis. We introduce $\textbf{Recon}$ ($\textbf{R}$easoning like an $\textbf{ECON}$omist), a 7B-parameter open-source LLM post-trained on a hand-curated dataset of 2,100 high-quality economic reasoning problems. Comprehensive evaluation on economic reasoning benchmarks and multi-agent games reveals clear improvements in structured reasoning and economic rationality. These results underscore the promise of domain-aligned post-training for enhancing reasoning and agent alignment, shedding light on the roles of SFT and RL in shaping model behavior. Code is available at https://github.com/MasterZhou1/Recon .

经济推理多智能体模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。