提出新方法提升长文本推理时缓存淘汰效率
Jacap: Robust KV Cache Eviction via Jacobian-Based Nonlinear Information Capacity Preservation

- 基于雅可比矩阵构建信息容量指标,量化查询相关性
- 在高压缩率下仍保持高精度,优于现有策略
- 适合需要高效长序列处理的模型部署场景
键值(KV)缓存淘汰对大语言模型的长上下文推理至关重要。然而,现有策略多依赖经验启发式方法,缺乏对软注意力机制下令牌价值的严格刻画。本文从局部信息几何角度重新思考该问题,将注意力过程建模为非线性高斯通信信道。通过注意力映射的一阶泰勒展开,推导出雅可比信息容量这一新目标,显式捕捉查询相关性、软注意敏感度与结构多样性。基于此理论,提出容量感知的淘汰方法 Jacap,采用软注意感知的重要性加权与统计杠杆率进行子集选择。在多种架构与基准上的大量实验表明,Jacap 在多数场景中表现更优,尤其在高压缩率条件下优势显著。
原文摘要 · Abstract (English)
Key-value (KV) cache eviction is essential for scaling long-context inference in Large Language Models. However, existing policies predominantly rely on empirical heuristics, lacking a rigorous characterization of token utility under the inherently nonlinear softmax attention mechanism. In this work, we rethink KV cache eviction through the lens of local information geometry, modeling the attention process as a nonlinear Gaussian communication channel. By performing a first-order Taylor expansion of the attention mapping, we derive the Jacobian Information Capacity, a novel objective that explicitly captures query relevance, softmax sensitivity, and structural diversity. Guided by this theory, we introduce Jacap, a capacity-aware eviction method that utilizes softmax-aware importance weighting and statistical leverage scores for subset selection. Extensive experiments across diverse architectures and benchmarks demonstrate that \textsc{Jacap} delivers superior performance in most scenarios, particularly in high-compression regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。