发现长文本中信息间距影响模型表现,提出新评测基准。
Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs
- 构建多信息片段间距评测集LongPiBench,模拟真实场景
- 6款开源与5款商用模型均在信息间隔大时性能下降
- 揭示间距而非位置是关键偏差来源,适合评估长文本模型
大型语言模型(LLMs)在处理长输入时存在位置偏差,尤其表现为“中间丢失”现象,即难以利用中间段落的相关信息。现有研究多关注单个相关片段,但实际应用常涉及多个信息点。为此,本文提出LongPiBench,一个评估多片段相关性位置偏差的基准。对五款商业模型和六款开源模型进行系统实验,结果表明:尽管多数模型对“中间丢失”具备一定鲁棒性,但在多个相关信息片段间距较大时仍存在显著偏差。该发现强调了评估与缓解位置偏差对提升长上下文能力的重要性。
原文摘要 · Abstract (English)
Positional bias in large language models (LLMs) hinders their ability to effectively process long inputs. A prominent example is the "lost in the middle" phenomenon, where LLMs struggle to utilize relevant information situated in the middle of the input. While prior research primarily focuses on single pieces of relevant information, real-world applications often involve multiple relevant information pieces. To bridge this gap, we present LongPiBench, a benchmark designed to assess positional bias involving multiple pieces of relevant information. Thorough experiments are conducted with five commercial and six open-source models. These experiments reveal that while most current models are robust against the "lost in the middle" issue, there exist significant biases related to the spacing of relevant information pieces. These findings highlight the importance of evaluating and reducing positional biases to advance LLM's capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。