构建巴西法律文本检索基准,统一评估标准。
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
- 设计统一评测框架,支持多类型法律文本检索
- 领域适配模型在特定子集上提升显著,BM25仍具竞争力
- 适合研究巴西法律信息检索的学者与开发者
葡萄牙语法律信息检索因现有数据集在文档类型、查询风格和相关性定义上差异大,难以系统评估。我们提出 JUÁ,一个面向巴西法律检索的公开基准,旨在支持跨异构法律语料的可复现、可比较评估。该基准不仅作为评测工具,更是一个持续评估基础设施,包含共享协议、统一排名指标、固定划分及公开排行榜。涵盖判例检索、立法、监管及问题驱动的法律搜索。我们评估了词法、稠密向量及基于 BM25 的重排序管道,包括在 JUÁ 对齐标注数据上微调的领域适配 Qwen 嵌入模型。结果表明,该基准具有足够多样性,可区分不同检索范式,并揭示显著的跨数据集权衡。领域适配在监督对齐的 JUÁ-Juris 子集上表现最佳,而 BM25 在存在强词汇与机构表述线索的场景中依然表现优异。总体而言,JUÁ 为在多个巴西法律领域下使用统一基准研究法律检索提供了实用框架。
原文摘要 · Abstract (English)
Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present JUÁ, a public benchmark for Brazilian legal retrieval designed to support more reproducible and comparable evaluation across heterogeneous legal collections. More broadly, JUÁ is intended not only as a benchmark, but as a continuous evaluation infrastructure for Brazilian legal IR, combining shared protocols, common ranking metrics, fixed splits when applicable, and a public leaderboard. The benchmark covers jurisprudence retrieval as well as broader legislative, regulatory, and question-driven legal search. We evaluate lexical, dense, and BM25-based reranking pipelines, including a domain-adapted Qwen embedding model fine-tuned on JUÁ-aligned supervision. Results show that the benchmark is sufficiently heterogeneous to distinguish retrieval paradigms and reveal substantial cross-dataset trade-offs. Domain adaptation yields its clearest gains on the supervision-aligned JUÁ-Juris subset, while BM25 remains highly competitive on other collections, especially in settings with strong lexical and institutional phrasing cues. Overall, JUÁ provides a practical evaluation framework for studying legal retrieval across multiple Brazilian legal domains under a common benchmark design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。