用新方法LeSeR提升法规问答中的文档检索精度
1-800-SHARED-TASKS at RegNLP: Lexical Reranking of Semantic Retrieval (LeSeR) for Regulatory Question Answering
- 提出词汇重排序方法LeSeR,优化语义检索排名
- 在10个召回中达到0.8201的召回率和0.6655的map
- 适合需要高精度法规信息检索的研究与应用
本文介绍了参加COLING 2025 RegNLP RIRAG挑战赛的系统方案,聚焦于监管领域中先进信息检索与答案生成技术的应用。我们测试了Stella、BGE、CDE和Mpnet等嵌入模型,结合微调与重排序策略,提升相关文档在前10位的检索效果。采用一种新方法LeSeR,取得了0.8201的recall@10和0.6655的map@10的优异表现。该工作展示了自然语言处理在监管场景中的潜力,为检索增强生成系统提供了实践依据,并指出了鲁棒性与领域适应性方面的改进方向。
原文摘要 · Abstract (English)
This paper presents the system description of our entry for the COLING 2025 RegNLP RIRAG (Regulatory Information Retrieval and Answer Generation) challenge, focusing on leveraging advanced information retrieval and answer generation techniques in regulatory domains. We experimented with a combination of embedding models, including Stella, BGE, CDE, and Mpnet, and leveraged fine-tuning and reranking for retrieving relevant documents in top ranks. We utilized a novel approach, LeSeR, which achieved competitive results with a recall@10 of 0.8201 and map@10 of 0.6655 for retrievals. This work highlights the transformative potential of natural language processing techniques in regulatory applications, offering insights into their capabilities for implementing a retrieval augmented generation system while identifying areas for future improvement in robustness and domain adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。