arXiv:2512.08078cs.IR2025-12

对比Azure AI搜索五种检索方法,助力法律团队高效发现关键信息

A Comparative Study of Retrieval Methods in Azure AI Search

  • 在RAG框架中测试关键词、语义、向量等五类检索策略
  • 混合检索法在准确率与相关性上表现最佳,优于单一方法
  • 适合法律从业者优化电子取证中的早期案件评估配置

越来越多的律师希望超越传统的关键词和语义搜索,以提升文档审查中关键信息的查找效率。大型语言模型(LLMs)现被视为律师可在文档审查中通过自然语言提问并获得精准简洁回答的工具。本研究评估了微软Azure检索增强生成(RAG)框架内的多种检索策略,旨在识别适用于电子发现(eDiscovery)中早期案件评估(ECA)的有效方法。在ECA阶段,法律团队需在正式大规模审查前对数据进行初步分析,以了解整体情况并识别关键事实与风险。本文对比了Azure AI Search的关键词、语义、向量、混合及混合-语义检索方法的性能,并呈现各方法生成答案的准确性、相关性与一致性。研究结果可帮助法律从业者未来更优地选择RAG配置。

原文摘要 · Abstract (English)

Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.

检索增强法律AIRAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。