arXiv:2412.15386cs.CLcs.AI2024-12EMNLP被引 6

测试大模型在长文本金融任务中的表现,发现越长越容易出错。

Systematic Evaluation of Long-Context LLMs on Financial Concepts

  • 用真实金融新闻构建渐进式任务评估长上下文模型
  • 长文本下性能急剧下降,简单任务也出现严重失效
  • 指令位置和格式微调都影响结果,适合严谨评测者参考

长上下文大语言模型(LC LLMs)有望提升现实场景中处理长文档的可靠性。然而,这些模型是否能真正有效利用不断扩大的上下文窗口仍待验证。本文通过构建真实世界金融新闻数据集,系统评估GPT-4系列先进LC LLMs在一系列逐步增加难度的任务上的表现,考察上下文长度、任务复杂度及关键信息位置的影响。结果表明,即使在简单任务中,长上下文模型在较长上下文时也表现出脆弱性,性能随任务复杂度上升而急剧下降。在长上下文条件下,这些先进模型出现灾难性指令遵循失败,导致退化输出。提示词消融实验还揭示模型对指令位置和微小标记格式仍高度敏感。最后,我们倡导采用综合指标如F1(而非仅召回率)并报告置信区间,以确保评估结果稳健且结论可靠。

原文摘要 · Abstract (English)

Long-context large language models (LC LLMs) promise to increase reliability of LLMs in real-world tasks requiring processing and understanding of long input documents. However, this ability of LC LLMs to reliably utilize their growing context windows remains under investigation. In this work, we evaluate the performance of state-of-the-art GPT-4 suite of LC LLMs in solving a series of progressively challenging tasks, as a function of factors such as context length, task difficulty, and position of key information by creating a real world financial news dataset. Our findings indicate that LC LLMs exhibit brittleness at longer context lengths even for simple tasks, with performance deteriorating sharply as task complexity increases. At longer context lengths, these state-of-the-art models experience catastrophic failures in instruction following resulting in degenerate outputs. Our prompt ablations also reveal unfortunate continued sensitivity to both the placement of the task instruction in the context window as well as minor markdown formatting. Finally, we advocate for more rigorous evaluation of LC LLMs by employing holistic metrics such as F1 (rather than recall) and reporting confidence intervals, thereby ensuring robust and conclusive findings.

长上下文金融文本模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。