提出新型长上下文评估框架,精准测试模型理解复杂结构的能力。
Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
- 通过隐式结构查询构建需剔除干扰信息的任务
- 在代码与自然语言任务中验证模型能力,显著优于传统检索测试
- 评估可自动打分,适合研究长上下文推理的模型
我们提出Michelangelo:一种简洁、合成且未泄露的长上下文推理评估方法,具备易于自动评分的特点。该评估基于全新的统一框架——隐式结构查询(Latent Structure Queries, LSQ),旨在衡量模型在长上下文中超越单一信息检索的能力。其核心思想是构造需要模型‘剔除’无关信息以揭示上下文潜在结构的任务,并通过询问结构细节来验证模型对隐含结构的理解。利用LSQ,我们在代码和自然语言领域构建了三个诊断性长上下文评估,旨在提供更强的模型能力信号。我们在多个先进模型上进行了测试,结果表明:(a) 所提评估具有高信号性;(b) 当前模型在合成长上下文信息方面仍有巨大提升空间。
原文摘要 · Abstract (English)
We introduce Michelangelo: a minimal, synthetic, and unleaked long-context reasoning evaluation for large language models which is also easy to automatically score. This evaluation is derived via a novel, unifying framework for evaluations over arbitrarily long contexts which measure the model's ability to do more than retrieve a single piece of information from its context. The central idea of the Latent Structure Queries framework (LSQ) is to construct tasks which require a model to ``chisel away'' the irrelevant information in the context, revealing a latent structure in the context. To verify a model's understanding of this latent structure, we query the model for details of the structure. Using LSQ, we produce three diagnostic long-context evaluations across code and natural-language domains intended to provide a stronger signal of long-context language model capabilities. We perform evaluations on several state-of-the-art models and demonstrate both that a) the proposed evaluations are high-signal and b) that there is significant room for improvement in synthesizing long-context information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。