arXiv:2411.07037cs.CL2024-11ACL被引 35

评测大模型在长文本中指令跟随的稳定性和能力

LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios

  • 构建可扩展数据集LIFBench,覆盖多种长文本场景和任务
  • 提出LIFEval评分方法,实现无需人工或模型辅助的自动评估
  • 测试20个主流大模型,揭示长上下文下性能差异与稳定性问题

随着大语言模型(LLMs)在自然语言处理中的发展,其在长上下文输入中稳定遵循指令的能力对实际应用至关重要。然而,现有基准测试很少关注长上下文中的指令跟随性或不同输入下的稳定性。为此,我们提出LIFBench,一个可扩展的数据集,用于评估LLM在长上下文中的指令跟随能力和稳定性。LIFBench包含三个长上下文场景和十一项多样化任务,涵盖2,766条通过自动化扩展方法生成的指令,涉及长度、表达方式和变量三个维度。为评估,我们提出LIFEval——一种基于评分标准的方法,可在不依赖人工判断或大模型辅助的情况下实现复杂响应的精确自动化打分。该方法支持从多角度全面分析模型性能与稳定性。我们在六个长度区间上对20个主流大模型进行了详细实验。本工作贡献了LIFBench和LIFEval,作为评估复杂长上下文设置下大模型表现的可靠工具,为未来模型发展提供关键洞察。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) evolve in natural language processing (NLP), their ability to stably follow instructions in long-context inputs has become critical for real-world applications. However, existing benchmarks seldom focus on instruction-following in long-context scenarios or stability on different inputs. To bridge this gap, we introduce LIFBench, a scalable dataset designed to evaluate LLMs' instruction-following capabilities and stability across long contexts. LIFBench comprises three long-context scenarios and eleven diverse tasks, featuring 2,766 instructions generated through an automated expansion method across three dimensions: length, expression, and variables. For evaluation, we propose LIFEval, a rubric-based assessment method that enables precise, automated scoring of complex LLM responses without reliance on LLM-assisted assessments or human judgment. This method allows for a comprehensive analysis of model performance and stability from multiple perspectives. We conduct detailed experiments on 20 prominent LLMs across six length intervals. Our work contributes LIFBench and LIFEval as robust tools for assessing LLM performance in complex and long-context settings, offering valuable insights to guide future advancements in LLM development.

大模型评测长上下文指令遵循自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。