arXiv:2409.06338cs.CL2024-09被引 3

区分长文本理解中检索与整体理解,量化任务类型与难度。

Retrieval Or Holistic Understanding? Dolce: Differentiate Our Long Context Evaluation Tasks

  • 用复杂度λ和冗余k参数化问题,划分五类关注点。
  • 通过采样短上下文估算模型解题概率,识别0%至67%为检索类任务。
  • 可区分模型是靠记忆还是理解,适合评估大模型长文本能力。

我们认为长文本理解包含两大核心能力:检索与整体理解。若不明确评估任务的聚焦类别,就无法真正理解和改进大模型的长文本能力。本文提出Dolce框架,通过λ(复杂度)和k(冗余度)参数化每个问题,并将其归入五个预定义的关注类别。我们建议从完整上下文中采样短片段,利用这些片段估计模型解决该问题的概率。为求得每个问题的λ和k,我们设计了一种混合模型,包含非参数背景噪声项与参数/非参数混合的基准组件,推导出在正确-错误(COW)和部分得分(PIG)两种场景下由λ和k参数化的概率函数。该方法在44个现有长文本评估任务中发现,0%至67%的问题属于检索聚焦,0%至90%的问题属于整体理解聚焦。

原文摘要 · Abstract (English)

We argue that there are two major distinct capabilities in long context understanding: retrieval and holistic understanding. Understanding and further improving LLMs' long context capabilities would not be possible without knowing the tasks' focus categories. We aim to automatically identify retrieval focused and holistic understanding focused problems from suites of benchmarks and quantitatively measure the difficulty within each focus. In this paper, we present the Dolce framework, which parameterizes each problem by $λ$ (complexity) and $k$ (redundancy) and assigns to one of five predefined focus categories. We propose to sample short contexts from the full context and estimate the probability an LLM solves the problem using the sampled spans. To find the $λ$ and $k$ for each problem, we further propose a mixture model of a non-parametric background noise component and a parametric/non-parametric hybrid oracle component, where we derive the probability functions parameterized by $λ$ and $k$ for both the correct-or-wrong (COW) scenario and the partial-point-in-grading (PIG) scenario. Our proposed methods can identify 0% to 67% of the problems are retrieval focused and 0% to 90% of the problems are holistic understanding focused across 44 existing long context evaluation tasks.

长文本理解评估框架大模型评测任务分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。