自一致性在长文本任务中会加剧位置偏见,反而降低模型表现。
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
- 在长上下文任务中,自一致性会放大位置偏差,导致性能下降。
- 模型越小、上下文越长,位置偏差越严重,且不受提示格式影响。
- 适用于需要全面理解长文档的场景,如法律或医学分析。
自一致性(SC)在短文本任务中能提升大语言模型性能,但在长上下文问题中是否依然有效?我们质疑该方法在长上下文场景下的普适性。由于位置偏见——模型对特定上下文区域的系统性依赖——使长文本理解受限。通过多种前沿模型、任务与SC实现方式的实验发现,自一致性不仅无法提升性能,反而显著恶化其表现。这种退化由持续存在的位置偏见驱动,在更长上下文和更小模型中加剧,但不受提示格式或任务类型影响。与短文本中分散推理路径不同,长文本下自一致性反而放大了位置错误。研究揭示了当前大模型在长上下文理解中的根本局限,亟需更先进的解决方案。
原文摘要 · Abstract (English)
Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC's benefits generalize to long-context settings, where LLMs often struggle with position bias, the systematic over-reliance on specific context regions-which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, we find that SC not only fails to improve but actively degrades performance on long-context tasks. This degradation is driven by persistent position bias, which worsens with longer context lengths and smaller model sizes but remains invariant to prompt format or task type. Unlike short-context tasks, where SC diversifies reasoning paths, long-context SC amplifies positional errors. These comprehensive results provide valuable insight into the limitations of current LLMs in long-context understanding and highlight the need for more sophisticated approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。