发现大模型在长文本处理中存在性能骤降临界点,影响实际应用。
Intelligence Degradation in Long-Context LLMs: Critical Threshold Determination via Natural Length Distribution Analysis
- 基于自然文本长度分析,排除人为干扰,定位性能下降根源。
- Qwen2.5-7B在上下文长度达最大值40%-50%时,F1分数从0.55降至0.3,降幅超45%
- 首次系统揭示开源模型长文本适应的局限性,指导实际部署
大语言模型在处理接近特定临界长度的上下文时,即使信息仍相关,也会出现灾难性性能下降,定义为任务性能下降超过30%,严重限制长文本应用。这种智能退化呈现普遍模式:模型在短至中等长度上下文中表现良好,但超过临界阈值后性能急剧崩溃。本文提出三种贡献:(1)自然长度分布分析:使用样本原始无截断、无填充的词元长度,提供更强因果证据,表明退化源于上下文长度本身;(2)临界阈值确定:在混合数据集(1,000个样本覆盖5%-95%上下文长度)上实验,确定Qwen2.5-7B在最大上下文长度40%-50%处存在临界点,此时F1分数从0.55-0.56下降至0.3,退化率达45.5%,通过五种方法交叉验证;(3)统一框架:整合浅层长上下文适应现象,解释退化模式并为缓解策略提供基础。本工作首次系统刻画了开源Qwen模型中的智能退化现象,为长上下文场景下的模型部署提供实用指导。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit catastrophic performance degradation when processing contexts approaching certain critical thresholds, even when information remains relevant. This intelligence degradation-defined as over 30% drop in task performance-severely limits long-context applications. This degradation shows a common pattern: models maintain strong performance up to a critical threshold, then collapse catastrophically. We term this shallow long-context adaptation-models adapt for short to medium contexts but fail beyond critical thresholds. This paper presents three contributions: (1) Natural Length Distribution Analysis: We use each sample's natural token length without truncation or padding, providing stronger causal evidence that degradation results from context length itself. (2) Critical Threshold Determination: Through experiments on a mixed dataset (1,000 samples covering 5%-95% of context length), we identify the critical threshold for Qwen2.5-7B at 40-50% of maximum context length, where F1 scores drop from 0.55-0.56 to 0.3 (45.5% degradation), using five-method cross-validation. (3) Unified Framework: We consolidate shallow adaptation, explaining degradation patterns and providing a foundation for mitigation strategies. This work provides the first systematic characterization of intelligence degradation in open-source Qwen models, offering practical guidance for deploying LLMs in long-context scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。