动态感知注意力与熵变,实现更高效的通用提示压缩
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
- 融合注意力机制与熵变化,动态调整压缩策略
- 在LongBench、GSM8K等数据集上显著提升压缩效果
- 适合需要长文本处理的大型语言模型应用
通用任务提示压缩通过利用自然语言中的冗余,降低计算开销并提升提示信息密度,尤其在长上下文场景中尤为重要。现有方法主要以信息熵为指标压缩词元,力求最小化信息损失。然而,这些方法忽略了两个关键问题:(i) 算法层面注意力关键词的重要性,(ii) 压缩过程中信息熵的变化。针对上述挑战,我们提出一种动态注意力感知的通用提示压缩方法(DAC)。该方法有效结合熵与注意力信息,动态感知压缩过程中的熵变,实现细粒度提示压缩。在LongBench、GSM8K和BBH等多个领域上的大量实验表明,DAC在多种任务和大模型上均表现出稳健且显著的改进,充分证明了其有效性。
原文摘要 · Abstract (English)
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in long-context scenarios. Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss. However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process. Motivated by these challenges, we propose a dynamic attention-aware approach for task-agnostic prompt compression (DAC). This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression. Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。