用合成数据提升越南法律文本检索准确率
Improving Vietnamese Legal Document Retrieval using Synthetic Data
- 用大模型生成多样化越南法律查询语句
- 预训练+难例挖掘使检索准确率显著提升
- 适合缺乏标注数据的低资源法律文本研究
在法律信息检索领域,基于嵌入的模型对精准问答系统至关重要。然而,大型标注数据集的匮乏构成重大挑战,尤其针对越南语法律文本。为此,我们提出一种新方法,利用大语言模型生成高质量、多样化的越南法律段落合成查询。这些合成数据用于预训练双编码器和ColBERT模型,并通过挖掘难负样本的对比损失进行微调。实验表明,这些改进显著提升了检索准确性,验证了合成数据与预训练技术在克服越南法律领域标注数据稀缺问题上的有效性。
原文摘要 · Abstract (English)
In the field of legal information retrieval, effective embedding-based models are essential for accurate question-answering systems. However, the scarcity of large annotated datasets poses a significant challenge, particularly for Vietnamese legal texts. To address this issue, we propose a novel approach that leverages large language models to generate high-quality, diverse synthetic queries for Vietnamese legal passages. This synthetic data is then used to pre-train retrieval models, specifically bi-encoder and ColBERT, which are further fine-tuned using contrastive loss with mined hard negatives. Our experiments demonstrate that these enhancements lead to strong improvement in retrieval accuracy, validating the effectiveness of synthetic data and pre-training techniques in overcoming the limitations posed by the lack of large labeled datasets in the Vietnamese legal domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。