测试大模型在真实长文本任务中的理解能力,发现其表现远低于宣称的上下文长度。
LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- 构建覆盖法律、金融等领域的自动化长文本基准数据集
- 6个本地部署模型最高仅达59.2%准确率,远低于预期
- 揭示当前大模型实际长依赖处理能力严重不足,适合关注实用场景的研究者
大语言模型虽具备日益扩展的上下文窗口,但在真实世界长依赖任务中的长期上下文理解能力仍存在根本性局限且研究不足。这一差距在法律、金融、游戏和代码等少被评测的真实应用场景中尤为显著。本文提出LooGLE v2,一个面向真实应用的新型基准,包含16k至200万词元的自动采集长文本,涵盖多个领域。我们设计了10类特定领域的长依赖任务,通过可扩展的数据整理流程生成1,934个具有多样性和复杂性的问答实例。对6个本地部署和4个API调用的LLM进行了全面评估,结果显示即使表现最佳的模型在该基准上也仅取得59.2%的整体得分。尽管具备广阔的上下文窗口,主流大模型的实际理解长度远小于其宣称能力,暴露出在真实长依赖任务中处理能力的巨大缺陷,凸显了模型在实际长上下文理解方面仍有巨大提升空间。
原文摘要 · Abstract (English)
Large language models (LLMs) are equipped with increasingly extended context windows recently, yet their long context understanding capabilities over long dependency tasks remain fundamentally limited and underexplored. This gap is especially significant in many real-world long-context applications that were rarely benchmarked. In this paper, we introduce LooGLE v2, a novel benchmark designed to evaluate LLMs' long context ability in real-world applications and scenarios. Our benchmark consists of automatically collected real-world long texts, ranging from 16k to 2M tokens, encompassing domains in law, finance, game and code. Accordingly, we delicately design 10 types of domain-specific long-dependency tasks and generate 1,934 QA instances with various diversity and complexity in a scalable data curation pipeline for further practical needs. We conduct a comprehensive assessment of 6 locally deployed and 4 API-based LLMs. The evaluation results show that even the best-performing model achieves only a 59.2% overall score on our benchmark. Despite the extensive context windows, popular LLMs are only capable of understanding a much shorter length of context than they claim to be, revealing significant limitations in their ability to handle real-world tasks with long dependencies and highlighting substantial room for model improvement in practical long-context understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。