用法律角色生成更丰富的法律查询数据,提升检索效果。
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
- 引入律师、检察官等角色生成多样化法律查询。
- 自相似度评分提升,原始语义保持不变。
- 适合法律信息检索、低资源领域数据增强场景。
低资源领域仍面临数据稀缺的挑战。现有数据增强方法虽利用大语言模型生成大量合成数据,但往往重数量轻质量,缺乏领域针对性。本文提出针对法律信息检索的基于角色的数据增强框架DALDALL,通过律师、检察官、法官等专业角色生成的合成查询,显著提升词汇与语义多样性。在CLERC和COLIEE基准上的实验表明,基于角色的增强在自相似度(Self-BLEU)上优于普通提示方法,同时保持原始查询的语义一致性。基于角色增强数据微调的稠密检索器,在召回率上表现优于或媲美原始数据训练模型。结果验证了角色提示是生成高质量专用数据的有效策略。
原文摘要 · Abstract (English)
Data scarcity remains a persistent challenge in low-resource domains. While existing data augmentation methods leverage the generative capabilities of large language models (LLMs) to produce large volumes of synthetic data, these approaches often prioritize quantity over quality and lack domain-specific strategies. In this work, we introduce DALDALL, a persona-based data augmentation framework tailored for legal information retrieval (IR). Our method employs domain-specific professional personas--such as attorneys, prosecutors, and judges--to generate synthetic queries that exhibit substantially greater lexical and semantic diversity than vanilla prompting approaches. Experiments on the CLERC and COLIEE benchmarks demonstrate that persona-based augmentation achieves improvement in lexical diversity as measured by Self-BLEU scores, while preserving semantic fidelity to the original queries. Furthermore, dense retrievers fine-tuned on persona-augmented data consistently achieve competitive or superior recall performance compared to those trained on original data or generic augmentations. These findings establish persona-based prompting as an effective strategy for generating high-quality training data in specialized, low-resource domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。