arXiv:2510.16783cs.CLcs.AI2025-10EMNLP

评测大模型在4千到12万词元长文本中的双语理解能力。

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

  • 构建中英双语长文本多任务评测集,涵盖四大挑战性任务。
  • 128k词元上下文下,即使GPT-4o也表现吃力,验证了难度。
  • 适合评估长文本推理、跨语言信息抽取等能力的模型开发者。

近年来,大语言模型在处理和理解长上下文方面展现出复杂能力,这要求更严格的评估方法来准确衡量其性能。本文提出 extbf{LC-Eval},一个面向英文与阿拉伯语的多任务长上下文理解评测基准,支持4k至超过128k词元的上下文长度。该基准包含四项新颖且具有挑战性的任务:多文档问答、双语问答、段落内主张验证,以及基于长上下文的多选题。这些任务旨在评估模型在深度推理、文档理解、信息追踪及双语信息提取与理解方面的能力。每个任务均提供中英文数据集,支持不同文本类型的跨语言对比分析。我们在开放权重与闭源大模型上进行了评估,结果显示,即便高性能模型如GPT-4o在部分任务上仍表现不佳,凸显了该基准的严苛性与实用性。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understanding. In this paper, we present \textbf{LC-Eval}, a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic, targeting context lengths ranging from 4k to over 128k tokens. LC-Eval introduces four novel and challenging tasks: multi-document question answering, bilingual question answering, claim verification within a paragraph, and multiple-choice questions based on long contexts. These tasks are designed to assess LLMs' abilities in deep reasoning, document comprehension, information tracing, and bilingual information extraction and understanding. The benchmark includes datasets in both Arabic and English for each task, allowing for a comparative analysis of their performance across different text genres. Evaluations were conducted on both open-weight and closed LLMs, with results indicating that LC-Eval presents significant challenges. Even high-performing models, such as GPT-4o, struggled with certain tasks, highlighting the complexity and rigor of the benchmark.

长文本理解多语言评测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。