arXiv:2506.08375cs.CL2025-06EMNLP被引 11

构建复杂指令遵循评测集,检验大模型多任务执行能力

EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models

  • 设计多任务并发+多种约束的复杂指令评测框架
  • 实测显示主流大模型在复杂指令下表现差异显著
  • 适合评估工业级应用中大模型的工作流执行能力

随着大语言模型(LLMs)的发展与广泛应用,'模型即产品'的新范式快速演进,对模型能力提出更高要求,尤其在需要精确工作流执行、准确理解多重任务的场景中。然而,现有基准大多聚焦于单一任务、约束有限的环境,难以真实反映现实应用复杂性。为此,我们提出极复杂指令遵循基准 EIFBENCH,旨在实现更真实、更鲁棒的 LLM 评估。EIFBENCH 不仅涵盖多任务并行场景,支持跨任务类型综合评估,还集成多种约束,模拟复杂操作环境。此外,我们提出分段策略优化(SegPO)算法,提升模型完成多任务工作流的准确性。在 EIFBENCH 上的评估揭示了现有 LLM 在面对极端复杂指令时存在显著性能差距,凸显了持续优化以应对大模型应用复杂挑战的必要性。

原文摘要 · Abstract (English)

With the development and widespread application of large language models (LLMs), the new paradigm of "Model as Product" is rapidly evolving, and demands higher capabilities to address complex user needs, often requiring precise workflow execution which involves the accurate understanding of multiple tasks. However, existing benchmarks focusing on single-task environments with limited constraints lack the complexity required to fully reflect real-world scenarios. To bridge this gap, we present the Extremely Complex Instruction Following Benchmark (EIFBENCH), meticulously crafted to facilitate a more realistic and robust evaluation of LLMs. EIFBENCH not only includes multi-task scenarios that enable comprehensive assessment across diverse task types concurrently, but also integrates a variety of constraints, replicating complex operational environments. Furthermore, we propose the Segment Policy Optimization (SegPO) algorithm to enhance the LLM's ability to accurately fulfill multi-task workflow. Evaluations on EIFBENCH have unveiled considerable performance discrepancies in existing LLMs when challenged with these extremely complex instructions. This finding underscores the necessity for ongoing optimization to navigate the intricate challenges posed by LLM applications.

指令遵循多任务评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。