用预期排名评估模型迁移能力,提升文本排序选型精度。
Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text Ranking
- 以预期排名作为迁移性度量,直接反映模型排序能力
- 自适应缩放句向量,缓解方向偏差并融合训练动态
- 在多个数据集上优于人工判断和现有方法,效率高
文本排序近年来因双编码器结合预训练语言模型(PLMs)取得显著进展。面对众多可用的PLMs,为特定数据集选择最优模型已成为一项挑战。尽管迁移性估计(TE)被视为替代人工直觉与暴力微调的有前景方案,但现有方法主要针对分类任务设计,其迁移性估计与文本排序目标不匹配。为此,本文提出将预期排名作为迁移性度量,显式反映模型的排序能力。同时,为缓解方向性偏差并融入训练动态,我们自适应地缩放各向同性的句子嵌入,以获得更准确的预期排名得分。所提方法Adaptive Ranking Transferability(AiRTran)能有效捕捉模型间的细微差异。在多个具有挑战性的文本排序数据集上,该方法显著优于以往面向分类的TE方法、人工直觉及ChatGPT,且耗时极少。
原文摘要 · Abstract (English)
Text ranking has witnessed significant advancements, attributed to the utilization of dual-encoder enhanced by Pre-trained Language Models (PLMs). Given the proliferation of available PLMs, selecting the most effective one for a given dataset has become a non-trivial challenge. As a promising alternative to human intuition and brute-force fine-tuning, Transferability Estimation (TE) has emerged as an effective approach to model selection. However, current TE methods are primarily designed for classification tasks, and their estimated transferability may not align well with the objectives of text ranking. To address this challenge, we propose to compute the expected rank as transferability, explicitly reflecting the model's ranking capability. Furthermore, to mitigate anisotropy and incorporate training dynamics, we adaptively scale isotropic sentence embeddings to yield an accurate expected rank score. Our resulting method, Adaptive Ranking Transferability (AiRTran), can effectively capture subtle differences between models. On challenging model selection scenarios across various text ranking datasets, it demonstrates significant improvements over previous classification-oriented TE methods, human intuition, and ChatGPT with minor time consumption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。