团队协作式框架提升多步骤任务表现
TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks
- 模拟人类团队分工,设四角色三阶段协作
- 在新基准上性能显著优于单模型
- 适合需要多步推理与角色协同的场景
近期多大语言模型框架被提出以解决上下文任务,但未显式模拟人类团队的角色分工,可能导致单一视角,削弱多步上下文任务的表现。为此,我们提出TeamLLM,一种类人团队导向的多大语言模型协作框架。TeamLLM采用四种不同角色进行明确分工,并通过三阶段协作机制处理多步上下文任务。为评估其在多步上下文任务中的有效性,我们构建了上下文锚定且流程结构化的任务(CGPST)基准,该基准具备四个核心特征:上下文锚定、流程结构、过程导向评估和多维评估。我们在整体、步骤和维度三个层面评估了十种主流大模型在CGPST上的表现。结果表明,TeamLLM在多项指标上显著提升性能。我们公开了包含场景、完整响应及人工评分的基准数据集。代码与数据可在https://anonymous.4open.science/r/TeamLLM-anonymous-C50E/ 获取。
原文摘要 · Abstract (English)
Recently, multi-Large Language Model (LLM) frameworks have been proposed to solve contextualized tasks. However, these frameworks do not explicitly emulate human team role division, which may lead to a single perspective, thereby weakening performance on multi-step contextualized tasks. To address this issue, we propose TeamLLM, a human-like Team-Oriented Multi-LLM Collaboration Framework. TeamLLM adopts four team roles with distinct division and employs a three-phase multi-LLM collaboration for multi-step contextualized tasks. To evaluate the effectiveness of TeamLLM on multi-step contextualized tasks, we propose Contextually-Grounded and Procedurally-Structured tasks (CGPST) and construct the CGPST benchmark. This benchmark has four core features: contextual grounding, procedural structure, process-oriented evaluation and multi-dimensional assessment. We evaluate ten popular LLMs on CGPST at overall-level, step-level, and dimension-level. Results show that TeamLLM substantially improves performance on CGPST. We release the benchmark with scenarios, full-process responses and human scores from ten LLMs. The code and data are available at https://anonymous.4open.science/r/TeamLLM-anonymous-C50E/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。