用位置间知识蒸馏缓解长文本中的位置偏差问题
Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- 通过优势位置向弱势位置蒸馏知识,缓解位置敏感性差异
- 在长文本检索与推理任务中实现各位置性能均匀提升
- 方法轻量高效,跨任务通用性强,适合长文本模型优化
位置偏差(PB)表现为不同上下文位置间敏感度不均,严重损害长文本理解与处理能力。以往方法或修改架构,或依赖大量上下文感知训练,前者无法消除显著性能差距,后者则带来巨大数据与计算开销。为此,我们提出 extbf{Pos2Distill}——一种位置到位置的知识蒸馏框架,利用优势位置向劣势位置转移能力,缩小性能差距。核心思想是借助内在的位置差异来对抗位置偏差本身。我们识别出在 extbf{ extsc{r}}etrieval与 extbf{ extsc{r}}easoning范式下PB的不同表现,分别设计了两个专用版本: extit{Pos2Distill-R extsuperscript{1}}和 extit{Pos2Distill-R extsuperscript{2}}。实验表明,该方法显著提升了所有上下文位置的性能均匀性,在长文本检索与推理任务中均取得显著性能增益。更重要的是,两者具备强跨任务泛化能力,且在各自任务上表现更优。
原文摘要 · Abstract (English)
Positional bias (PB), manifesting as non-uniform sensitivity across different contextual locations, significantly impairs long-context comprehension and processing capabilities. Previous studies have addressed PB either by modifying the underlying architectures or by employing extensive contextual awareness training. However, the former approach fails to effectively eliminate the substantial performance disparities, while the latter imposes significant data and computational overhead. To address PB effectively, we introduce \textbf{Pos2Distill}, a position to position knowledge distillation framework. Pos2Distill transfers the superior capabilities from advantageous positions to less favorable ones, thereby reducing the huge performance gaps. The conceptual principle is to leverage the inherent, position-induced disparity to counteract the PB itself. We identify distinct manifestations of PB under \textbf{\textsc{r}}etrieval and \textbf{\textsc{r}}easoning paradigms, thereby designing two specialized instantiations: \emph{Pos2Distill-R\textsuperscript{1}} and \emph{Pos2Distill-R\textsuperscript{2}} respectively, both grounded in this core principle. By employing the Pos2Distill approach, we achieve enhanced uniformity and significant performance gains across all contextual positions in long-context retrieval and reasoning tasks. Crucially, both specialized systems exhibit strong cross-task generalization mutually, while achieving superior performance on their respective tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。