arXiv:2409.01357cs.CLcs.IR2024-09被引 2

研究法语法律领域混合检索,发现零样本下融合有效,但领域内训练后需调权才不拖累性能。

Know When to Fuse: Investigating Non-English Hybrid Retrieval in the Legal Domain

  • 对比多种主流检索模型,测试其在法语法律文本中的混合效果。
  • 零样本时融合提升性能,领域内训练后融合常降低效果。
  • 揭示非英语专业领域中混合检索的适用边界,适合法律信息检索研究者。

混合检索已成为克服不同匹配范式局限性的有效策略,尤其在跨领域场景中显著提升了检索质量。然而,现有研究主要聚焦有限的检索方法,且仅在英文通用数据集上以成对方式评估。本文首次在法语法律领域探索多种主流检索模型的混合效果,涵盖零样本与领域内两种情景。结果表明,在零样本条件下,融合不同通用模型始终优于单一模型,无论采用何种融合方法;而当模型在领域内训练后,融合通常会降低性能,除非对得分进行精细加权调整。这些新发现拓展了既有结论在新领域和语言中的适用性,深化了对非英语专业领域混合检索机制的理解。

原文摘要 · Abstract (English)

Hybrid search has emerged as an effective strategy to offset the limitations of different matching paradigms, especially in out-of-domain contexts where notable improvements in retrieval quality have been observed. However, existing research predominantly focuses on a limited set of retrieval methods, evaluated in pairs on domain-general datasets exclusively in English. In this work, we study the efficacy of hybrid search across a variety of prominent retrieval models within the unexplored field of law in the French language, assessing both zero-shot and in-domain scenarios. Our findings reveal that in a zero-shot context, fusing different domain-general models consistently enhances performance compared to using a standalone model, regardless of the fusion method. Surprisingly, when models are trained in-domain, we find that fusion generally diminishes performance relative to using the best single system, unless fusing scores with carefully tuned weights. These novel insights, among others, expand the applicability of prior findings across a new field and language, and contribute to a deeper understanding of hybrid search in non-English specialized domains.

法律AI混合检索法语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。