自然文本演变会严重削弱大模型的问答能力,即使问题和信息不变。
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- 构建人类编辑的段落变体框架,量化语义相似度
- 相似度越低,模型准确率下降超30%,部分模型斜率超70
- 适合关注模型泛化与长期可用性的研究者
大语言模型在自然演化文本上的理解能力如何?为此,我们提出一种框架,用于构建来自当代问答基准的自然演化的段落变体,并分析模型在不同语义相似度下的表现。该相似度衡量变体与预训练时所见内容的接近程度。我们在六个问答数据集和八种公开训练的大语言模型上进行了评估。结果表明,当段落自然偏离预训练版本时,即使问题和必要信息在推理时仍存在,模型性能也显著下降。例如,在BoolQ数据集上,平均准确率从高相似度到低相似度区间下降超过30%,多个模型的下降斜率超过70。这说明自然文本演化对大模型的语言理解能力构成重大挑战。
原文摘要 · Abstract (English)
How does the natural evolution of context paragraphs affect question answering in generative Large Language Models (LLMs)? To investigate this, we propose a framework for curating naturally evolved, human-edited variants of reading passages from contemporary QA benchmarks and for analyzing LLM performance across a range of semantic similarity scores, which quantify how closely each variant aligns with content seen during pretraining. Using this framework, we evaluate six QA datasets and eight LLMs with publicly available training data. Our experiments reveal that LLM performance declines as reading passages naturally diverge from the versions encountered during pretraining-even when the question and all necessary information remains present at inference time. For instance, average model accuracy on BoolQ drops by over 30% from the highest to lowest similarity bins, with slopes exceeding 70 across several LLMs. These findings suggest that natural text evolution poses a significant challenge to the language understanding capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。