arXiv:2601.23228cs.AIcs.CL2026-01被引 1

用过程奖励提升多智能体系统训练效率,无需人工标注。

Scaling Multiagent Systems with Process Rewards

  • 为每个动作分配AI反馈的奖励,实现细粒度监督。
  • 数学竞赛任务上准确率提升5.0至17.5个百分点。
  • 适合长时序、复杂任务的多智能体系统研究者。

多智能体系统在通过分工解决复杂任务方面展现出潜力,但同时微调多个智能体面临两大挑战:(1)跨智能体的信用分配问题;(2)昂贵的多智能体轨迹采样效率低。本文提出基于AI反馈的逐动作过程奖励(MAPPA)来解决这两个问题。通过将信用分配给单个智能体的动作而非仅在任务完成时,MAPPA在无需真实标签的情况下实现细粒度监督,并从每次轨迹中提取最大训练信号。我们在竞赛数学题和工具增强的数据分析任务上验证该方法。在未见的数学问题上,MAPPA在AIME上提升5.0–17.5个百分点,在AMC上提升7.8–17.2个百分点。在数据分析任务中,成功率提升16.7个百分点,质量指标最高提升47%。结果表明,逐动作监督可在不同多智能体系统和领域中带来显著改进。本工作为实现少人工干预的复杂、长时序多智能体系统奠定了初步基础。

原文摘要 · Abstract (English)

While multiagent systems have shown promise for tackling complex tasks via specialization, finetuning multiple agents simultaneously faces two key challenges: (1) credit assignment across agents, and (2) sample efficiency of expensive multiagent rollouts. In this work, we propose finetuning multiagent systems with per-action process rewards from AI feedback (MAPPA) to address both. Through assigning credit to individual agent actions rather than only at task completion, MAPPA enables fine-grained supervision without ground truth labels while extracting maximal training signal from each rollout. We demonstrate our approach on competition math problems and tool-augmented data analysis tasks. On unseen math problems, MAPPA achieves +5.0--17.5pp on AIME and +7.8--17.2pp on AMC. For data analysis tasks, our method improves success rate by +16.7pp while quality metrics improve by up to 47%, validating that per-action supervision can lead to improvements across different multiagent systems on various domains. By addressing these challenges, our work takes a first step toward scaling multiagent systems for complex, long-horizon tasks with minimal human supervision.

多智能体过程奖励自动标注数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。