arXiv:2412.19265cs.IRcs.CL2024-12

针对日本法律文本优化检索模型,提升准确率与效率

Optimizing Multi-Stage Language Models for Effective Text Retrieval

  • 分两阶段设计检索流程,融合多策略集成增强适应性
  • 在日文法律数据集上表现超越现有方法,MS-MARCO也显著提升
  • 适合法律、多语言场景下的复杂查询任务

高效文本检索对法律文档分析等应用至关重要,尤其在日语法律系统等专业领域。现有方法在特定领域表现不佳,亟需定制化方案。本文提出一种针对日语法律数据集的新型两阶段检索管道,利用先进语言模型实现当前最优性能,显著提升检索效率与准确率。通过集成多种检索策略的组合模型,进一步增强鲁棒性与适应性,在多样任务中取得更优结果。大量实验验证了该方法的有效性,不仅在日文法律数据集上表现突出,也在广泛使用的基准数据集MS-MARCO上展现强劲性能。本工作为领域特定及通用场景下的文本检索设立了新标准,提供解决法律与多语言环境下复杂查询的综合性方案。

原文摘要 · Abstract (English)

Efficient text retrieval is critical for applications such as legal document analysis, particularly in specialized contexts like Japanese legal systems. Existing retrieval methods often underperform in such domain-specific scenarios, necessitating tailored approaches. In this paper, we introduce a novel two-phase text retrieval pipeline optimized for Japanese legal datasets. Our method leverages advanced language models to achieve state-of-the-art performance, significantly improving retrieval efficiency and accuracy. To further enhance robustness and adaptability, we incorporate an ensemble model that integrates multiple retrieval strategies, resulting in superior outcomes across diverse tasks. Extensive experiments validate the effectiveness of our approach, demonstrating strong performance on both Japanese legal datasets and widely recognized benchmarks like MS-MARCO. Our work establishes new standards for text retrieval in domain-specific and general contexts, providing a comprehensive solution for addressing complex queries in legal and multilingual environments.

文本检索法律AI多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。