arXiv:2608.20099cs.MAcs.CL2026-08中稿 · ICONIP 2026

用奖励机制引导生成更简洁的多智能体通信拓扑,降低计算开销。

Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design

论文配图:Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design
图 1 · 摘自论文原文
  • 引入奖励模型,同时优化任务正确性和结构紧凑性
  • 在保持原有准确率基础上,平均减少20.5%的令牌消耗
  • 适合需要高效通信的多智能体系统设计场景

基于大语言模型的多智能体系统(MAS)在复杂推理任务中表现优异,但需消耗大量令牌。现有自动拓扑设计方法ARG-Designer将问题建模为自回归图生成,但其训练目标未显式鼓励生成稀疏高效的拓扑。为此,我们提出受奖励引导的自回归图生成方法(RGA-Designer),灵感来自人类反馈强化学习(RLHF)。通过联合捕捉任务正确性与结构紧凑性的奖励模型,对预训练图生成器进行微调。该方法在保持ARG-Designer任务准确率的同时,平均降低20.5%的令牌消耗。

原文摘要 · Abstract (English)

LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption. Recent work on automatic topology design, ARG-Designer, has reframed this problem as autoregressive graph generation. However, its training objective provides no explicit incentive for the model to generate sparse and efficient topologies. We address this limitation by introducing a Reward-Guided Autoregressive Graph Generation (RGA-Designer) inspired by Reinforcement Learning from Human Feedback (RLHF). We train a reward model that jointly captures task correctness and structural compactness, and then fine-tune the pretrained graph generator using the reward model as feedback. Our method preserves task accuracy at the level of ARG-Designer while reducing token consumption by an average of 20.5%.

多智能体图生成奖励机制效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。