arXiv:2505.14925cs.CLcs.AI2025-05被引 10

用小说测试大模型长文本理解,发现64k后能力大幅下降

Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels

  • 用小说构建新评测基准,考察剧情、世界观和时间线理解
  • 7个前沿大模型在超过64k token后表现显著下滑
  • 适合关注长文本推理的开发者和研究者使用

尽管大语言模型(LLMs)的上下文长度已扩展至百万级别,但评估其在超越‘大海捞针’式任务之外的有效性仍具挑战。我们提出,小说是具有细微复杂结构与长程语义依赖的典型场景,常长达128k token以上。受计算小说分析研究启发,我们发布太长没模型(TLDM)基准,用于测试模型报告情节摘要、故事世界配置及叙事时间流逝的能力。实验发现,七种前沿大模型在超过64k token后均失去稳定理解能力。结果表明,模型开发者需超越‘中间丢失’类基准,以评估复杂长上下文场景下的真实性能。为促进后续研究,我们公开了TLDM基准、参考代码与数据。

原文摘要 · Abstract (English)

Although the context length of large language models (LLMs) has increased to millions of tokens, evaluating their effectiveness beyond needle-in-a-haystack approaches has proven difficult. We argue that novels provide a case study of subtle, complicated structure and long-range semantic dependencies often over 128k tokens in length. Inspired by work on computational novel analysis, we release the Too Long, Didn't Model (TLDM) benchmark, which tests a model's ability to report plot summary, storyworld configuration, and elapsed narrative time. We find that none of seven tested frontier LLMs retain stable understanding beyond 64k tokens. Our results suggest language model developers must look beyond "lost in the middle" benchmarks when evaluating model performance in complex long-context scenarios. To aid in further development we release the TLDM benchmark together with reference code and data.

长文本理解大模型评测小说分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。