arXiv:2608.20853cs.AI2026-08

首个多语言细粒度长文本评测基准,覆盖六语种文档级理解

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

论文配图:MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
图 1 · 摘自论文原文
  • 基于联合国报告构建跨语言细粒度长文本数据集,按词句段文档分层
  • 模型在词级别表现好,但在段落级任务中显著下降,低资源语言差距大
  • 揭示模型依赖表面连接词、生成流畅但事实不一致等新问题

长上下文大语言模型的评估进展迅速,但现有基准多局限于文档级且集中于高资源语言,未能充分覆盖细粒度挑战。为此,我们提出MGAL——首个多语言、细粒度与位置感知的长上下文评测基准。该基准基于涵盖8K至128K token的联合国报告,覆盖六种官方语言,包含词、句、段、文档四个语言粒度层次,并按内容在文档中的位置(开头、中间、结尾)及段落层级进行分层索引。该设计支持对多语言长文本理解能力的系统性诊断。实验发现:(1) 模型在词级任务表现良好,但在更粗粒度任务中表现下降;(2) 封闭源模型在低资源语言中仍具明显优势。进一步识别出两个新挑战:(1) 局部语义拥挤下,模型更依赖表面线索(如‘however’等连接词或重复实体),而非句子在上下文中的语篇角色(如背景、结果);(2) 生成文本流畅性与一致性之间存在差距,即输出读起来顺滑但偏离原文事实。此外,观察到若干与前期研究一致的现象,包括对邻近证据的依赖和不确定时重复选项的倾向。

原文摘要 · Abstract (English)

Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like ``however'' or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.

长文本理解多语言评测基准细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。