arXiv:2506.02921cs.CL2025-06NeurIPS被引 7

用人工传记构建可控评估框架,更真实检验长文本模型能力

A Controllable Examination for Long-Context Language Models

  • 以生成传记为背景,实现上下文连贯的可控测试环境
  • 18个模型测试显示,长文本下理解与推理能力普遍不足
  • 相比旧基准,更贴近真实任务且结果可解释,适合研究者使用

现有长文本语言模型(LCLM)评估框架可分为真实应用(如文档摘要)和合成任务(如藏针于海)。前者复杂难解且易数据污染,后者常因目标信息与上下文缺乏语义关联而削弱有效性。为此,我们提出理想评估应具备:1)连贯上下文;2)可控设置;3)可靠评估。本文引入$ extbf{LongBioBench}$,利用人工生成的传记作为受控环境,评估模型在理解、推理与可信度方面的表现。共测试18个LCLM,结果表明多数模型在长上下文中仍存在语义理解与基础推理缺陷,且可信度随上下文长度增加而下降。进一步分析发现,现有合成基准中存在的非连贯上下文、数值型目标、缺少干扰项等问题,使其难以有效检验模型能力。相较之下,LongBioBench在模拟真实任务与保持可控性之间取得更好平衡,具备高度可解释性与可配置性。

原文摘要 · Abstract (English)

Existing frameworks for evaluating long-context language models (LCLM) can be broadly categorized into real-world applications (e.g, document summarization) and synthetic tasks (e.g, needle-in-a-haystack). Despite their utility, both approaches are accompanied by certain intrinsic limitations. Real-world tasks often involve complexity that makes interpretation challenging and suffer from data contamination, whereas synthetic tasks frequently lack meaningful coherence between the target information (needle) and its surrounding context (haystack), undermining their validity as proxies for realistic applications. In response to these challenges, we posit that an ideal long-context evaluation framework should be characterized by three essential features: 1) seamless context 2) controllable setting and 3) sound evaluation. This study introduces $\textbf{LongBioBench}$, a benchmark that utilizes artificially generated biographies as a controlled environment for assessing LCLMs across dimensions of understanding, reasoning, and trustworthiness. Our experimental evaluation, which includes 18 LCLMs in total, demonstrates that most models still exhibit deficiencies in semantic understanding and elementary reasoning over retrieved results and are less trustworthy as context length increases. Our further analysis indicates some design choices employed by existing synthetic benchmarks, such as contextual non-coherence, numerical needles, and the absence of distractors, rendering them vulnerable to test the model's long-context capabilities. To sum up, compared to previous synthetic benchmarks, LongBioBench achieves a better trade-off between mirroring authentic language tasks and maintaining controllability, and is highly interpretable and configurable.

长文本评估可控实验语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。