arXiv:2604.05253cs.IRcs.LG2026-04中稿 · the 1st Late Inter…

发现晚交互检索中最大相似度聚合会引发梯度集中,影响模型稳定性。

Spike Hijacking in Late-Interaction Retrieval

  • 用合成数据揭示最大相似度聚合导致梯度集中在少数片段。
  • 文档越长,最大相似度方法性能下降越快,优于平滑池化的方案更鲁棒。
  • 适合关注多向量检索系统设计与梯度稳定性研究的读者。

晚交互检索模型依赖硬最大相似度(MaxSim)聚合词元级相似性。尽管有效,这种‘赢家通吃’的池化规则可能结构性地扭曲训练动态。我们在一个受控的批次对比训练合成环境中,证明了MaxSim比Top-k池化和softmax聚合引发显著更高的块级梯度集中。虽然稀疏路由有助于早期区分,但也增加对文档长度的敏感性:随着文档块数量增加,MaxSim的性能退化远超温和平滑变体。我们在真实世界的多向量检索基准上验证了这些发现,控制文档长度的实验揭示了硬最大池化下的类似脆弱性。结果表明,池化引起的梯度集中是晚交互检索的结构性特征,并凸显了稀疏性与鲁棒性之间的权衡。这为多向量检索系统中硬最大池化的替代方案提供了理论依据。

原文摘要 · Abstract (English)

Late-interaction retrieval models rely on hard maximum similarity (MaxSim) to aggregate token-level similarities. Although effective, this winner-take-all pooling rule may structurally bias training dynamics. We provide a mechanistic study of gradient routing and robustness in MaxSim-based retrieval. In a controlled synthetic environment with in-batch contrastive training, we demonstrate that MaxSim induces significantly higher patch-level gradient concentration than smoother alternatives such as Top-k pooling and softmax aggregation. While sparse routing can improve early discrimination, it also increases sensitivity to document length: as the number of document patches grows, MaxSim degrades more sharply than mild smoothing variants. We corroborate these findings on a real-world multi-vector retrieval benchmark, where controlled document-length sweeps reveal similar brittleness under hard max pooling. Together, our results isolate pooling-induced gradient concentration as a structural property of late-interaction retrieval and highlight a sparsity-robustness tradeoff. These findings motivate principled alternatives to hard max pooling in multi-vector retrieval systems.

检索模型梯度集中多向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。