研究发现长文本生成中事实准确性随长度下降,主因是知识耗尽。
How Does Response Length Affect Long-Form Factuality
- 提出双层级自动化评估框架,高效且贴近人工判断
- 实验显示响应越长,事实精确率越低,存在长度偏差
- 发现知识耗尽是事实错误主因,非错误传播或上下文过长
大型语言模型广泛用于长文本生成,但回答中的事实错误会削弱其可靠性。尽管学界日益关注模型事实性,响应长度对事实性的影响仍研究不足。本文首次系统探究该关系,提出一种自动化的双层级长文本事实性评估框架,在保持高人工标注一致性的同时具备成本优势。基于此框架,我们开展受控实验,发现更长的响应表现出更低的事实精确率,证实了长度偏差的存在。为解释这一现象,我们实证检验了三种假设:错误传播、长上下文和知识耗尽。结果表明,知识耗尽(模型逐渐耗尽可靠知识)是事实退化的主要原因,而非其余两个假设。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used for long-form text generation. However, factual errors in the responses would undermine their reliability. Despite growing attention to LLM factuality, the effect of response length on factuality remains underexplored. In this work, we systematically investigate this relationship by first introducing an automatic and bi-level long-form factuality evaluation framework, which achieves high agreement with human annotations while being cost-effective. Using this framework, we conduct controlled experiments and find that longer responses exhibit lower factual precision, confirming the presence of length bias. To explain this phenomenon, we empirically examine three hypotheses: error propagation, long context, and facts exhaustion. Our results reveal that facts exhaustion, where the model gradually exhausts more reliable knowledge, is the primary cause of factual degradation, rather than the other two hypotheses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。