让大模型通过多人互动学习更强推理能力
MARO: Learning Stronger Reasoning from Social Interaction
- 将结果拆解为交互中的具体行为,解决奖励信号稀疏问题
- 平衡不同角色训练样本权重,缓解角色分布不均
- 直接评估行为效用,提升环境稳定性,适合研究通用推理
人类在日常生活中面临无数需要推理与判断的场景。然而,现有大语言模型训练方法主要依赖文本内容或预设问题,缺乏与他人交互、协商和竞争的真实经验。为此,本文提出多智能体奖励优化(MARO),使大语言模型通过多智能体社会环境的学习与实践,获得更强的推理能力。具体而言,MARO首先通过将最终成败结果分解为交互过程中的每个具体行为,解决奖励信号稀疏问题;其次,通过平衡不同角色的训练样本权重,缓解角色分布不均问题;最后,通过直接评估每个行为的效用,应对环境不稳定问题。实验表明,MARO不仅显著提升模型的社会推理能力,且通过社会模拟学习获得的能力可有效迁移到数学推理和指令遵循等任务中。这揭示了多智能体社会学习在增强大语言模型通用推理能力方面的巨大潜力。
原文摘要 · Abstract (English)
Humans face countless scenarios that require reasoning and judgment in daily life. However, existing large language model training methods primarily allow models to learn from existing textual content or solve predetermined problems, lacking experience in real scenarios involving interaction, negotiation, and competition with others. To address this, this paper proposes Multi-Agent Reward Optimization (MARO), a method that enables large language models (LLMs) to acquire stronger reasoning abilities by learning and practicing in multi-agent social environments. Specifically, MARO first addresses the sparse learning signal problem by decomposing final success or failure outcomes into each specific behavior during the interaction process; second, it handles the uneven role distribution problem by balancing the training sample weights of different roles; finally, it addresses environmental instability issues by directly evaluating the utility of each behavior. Experimental results demonstrate that MARO not only achieves significant improvements in social reasoning capabilities, but also that the abilities acquired through social simulation learning can effectively transfer to other tasks such as mathematical reasoning and instruction following. This reveals the tremendous potential of multi-agent social learning in enhancing the general reasoning capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。