arXiv:2601.04395cs.IR2026-01ACL

调整相关性阈值可显著提升多语言稠密检索效果

The Overlooked Role of Graded Relevance Thresholds in Multilingual Dense Retrieval

  • 用分级相关性分数替代二元判断,更贴近真实反馈
  • 合理设置阈值可提升检索效果并减少所需标注数据
  • 适合关注多语言检索优化的研究者和工程师

稠密检索模型通常通过对比学习进行微调,要求使用二元相关性判断,尽管相关性本质上是分级的。本文分析了分级相关性分数及用于转换为二元标签的阈值对多语言稠密检索的影响。基于大模型标注的相关性评分构建的多语言数据集,研究了单语、混合语种和跨语言检索场景。结果表明,最优阈值在不同语言和任务间系统性变化,常与资源水平差异相关。恰当的阈值可提升性能、减少微调数据量,并缓解标注噪声;而错误选择则会降低效果。我们认为分级相关性是被忽视的重要信号,阈值校准应作为微调流程的规范环节。

原文摘要 · Abstract (English)

Dense retrieval models are typically fine-tuned with contrastive learning objectives that require binary relevance judgments, even though relevance is inherently graded. We analyze how graded relevance scores and the threshold used to convert them into binary labels affect multilingual dense retrieval. Using a multilingual dataset with LLM-annotated relevance scores, we examine monolingual, multilingual mixture, and cross-lingual retrieval scenarios. Our findings show that the optimal threshold varies systematically across languages and tasks, often reflecting differences in resource level. A well-chosen threshold can improve effectiveness, reduce the amount of fine-tuning data required, and mitigate annotation noise, whereas a poorly chosen one can degrade performance. We argue that graded relevance is a valuable but underutilized signal for dense retrieval, and that threshold calibration should be treated as a principled component of the fine-tuning pipeline.

稠密检索多语言相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。