arXiv:2601.18527cs.CL2026-01Conference of the …

优化长文本模型的微调策略,提升信息检索与缓存压缩下的性能。

Exploring Fine-Tuning for In-Context Retrieval and Efficient KV-Caching in Long-Context Language Models

  • 设计微调方法增强模型在长上下文中的相关信息识别能力。
  • 域内任务最高提升20分,金融类问题增益达+9分。
  • 对缓存压缩有中等抗扰性,适用于需要高效推理的场景。

随着上下文窗口扩展至数百万标记,长上下文语言模型(LCLMs)能够编码完整文档集合,为传统检索增强生成(RAG)提供有力替代方案。然而,微调策略是否能有效提升长上下文表现,以及在键值缓存压缩下的鲁棒性仍不明确。本文研究不同训练策略对LCLMs定位并利用相关信息能力的影响,以及其在KV缓存压缩下的稳健性。实验显示,域内任务性能显著提升,最高达+20分;金融类问题表现突出,较基线提升+9分,而多选题任务中RAG仍优于基线+6分。此外,微调策略在缓存压缩下带来适度鲁棒性提升,效果因任务而异。

原文摘要 · Abstract (English)

With context windows of millions of tokens, Long-Context Language Models (LCLMs) can encode entire document collections, offering a strong alternative to conventional retrieval-augmented generation (RAG). However, it remains unclear whether fine-tuning strategies can improve long-context performance and translate to greater robustness under KV-cache compression techniques. In this work, we investigate which training strategies most effectively enhance LCLMs' ability to identify and use relevant information, as well as enhancing their robustness under KV-cache compression. Our experiments show substantial in-domain improvements, achieving gains of up to +20 points over the base model. However, out-of-domain generalization remains task dependent with large variance -- LCLMs excels on finance questions (+9 points), while RAG shows stronger performance on multiple-choice questions (+6 points) over the baseline models. Finally, we show that our fine-tuning approaches bring moderate improvements in robustness under KV-cache compression, with gains varying across tasks.

长文本建模微调KV缓存

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。