arXiv:2509.17628cs.CLcs.AI2025-09被引 4

构建多阶段协作推理新基准,评估大模型在复杂场景中的协同能力。

MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents

  • 设计三阶段流水线生成高质量领域问答数据
  • 涵盖4大领域12.6万条数据,分三级难度测试复杂推理
  • 揭示商业模型仍存复杂任务差距,且对噪声敏感

大型语言模型在单一领域问答任务中表现优异,但在复杂多阶段场景中的推理与协作能力仍待探索。现有基准多聚焦孤立任务或窄领域,忽视模型在无外部引导下的多阶段协同与优化能力。为此,我们提出MSCoRe,一个包含126,696条领域特定问答实例的新基准,覆盖汽车、医药、电子和能源领域。数据通过动态采样、迭代问答生成与多级质量评估的三阶段流程构建,任务按阶段覆盖度和复杂度分为三级难度。我们对多种前沿大模型代理进行了全面评估:商用模型整体表现最佳,但简单与复杂任务间仍存在显著的ROUGE分数差距;同时发现模型性能受噪声数据负面影响。MSCoRe为社区提供了评估和提升大模型多阶段推理能力的重要资源。代码与数据见https://github.com/D3E0-source/MSCoRE。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have excelled in question-answering (QA) tasks within single domains. However, their reasoning and coordination capabilities in complex, multi-stage scenarios remain underexplored. Existing benchmarks typically focus on isolated tasks or narrow domains, overlooking models' abilities for multi-stage collaboration and optimization without explicit external guidance. To bridge this gap, we propose \textbf{MSCoRe}, a novel benchmark comprising 126696 domain-specific QA instances spanning scenarios in automotive, pharmaceutical, electronics, and energy sectors. The dataset is created using a structured three-phase pipeline: dynamic sampling, iterative question-answer generation, and a multi-level quality assessment to ensure data quality. Tasks are further categorized into three difficulty levels according to stage coverage and complexity. With MSCoRe, we have conducted a comprehensive evaluation of various state-of-the-art LLM agents. The commercial models performed best across all tasks and scenarios, but a notable gap in ROUGE scores remains between simple and complex tasks. We also tested the models' robustness and found that their performance is negatively affected by noisy data. MSCoRe provides a valuable new resource for the community to evaluate and improve multi-stage reasoning in LLM agents. The code and data are available at https://github.com/D3E0-source/MSCoRE.

大模型推理多阶段协作评测基准LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。