arXiv:2604.00731cs.IR2026-04被引 1

用自动化方法生成阿尔及利亚法律检索测试集,省去99%人工标注。

STCALIR: Semi-Synthetic Test Collection for Algerian Legal Information Retrieval

  • 从原始法律文档自动构建半合成测试集,减少人工标注。
  • 检索效果达人类标注水平(Hit@10≈0.785),相关性评估高度一致。
  • 适合低资源法律领域研究者快速搭建可复现评估体系。

测试集对检索与重排序模型的评估至关重要。然而,在阿尔及利亚法律等专业领域,由于高质量语料和相关性判断稀缺,构建测试集成本高昂。为此,我们提出STCALIR框架,直接从原始法律文档生成半合成测试集。该流程遵循Cranfield范式,保留主题、语料库和相关性判断三大核心组件,通过自动化多阶段检索与过滤,将标注工作量减少99%。我们在Mr. TyDi基准上验证了STCALIR,结果表明生成的半合成相关性判断在检索效果上与人工标注相当(Hit@10 ≈ 0.785)。此外,基于这些标签的系统级排序与人工评估高度一致,肯德尔τ为0.89,斯皮尔曼ρ为0.92。总体而言,STCALIR为低资源法律领域提供了一种可复现且低成本的可靠测试集构建方案。

原文摘要 · Abstract (English)

Test collections are essential for evaluating retrieval and re-ranking models. However, constructing such collections is challenging due to the high cost of manual annotation, particularly in specialized domains like Algerian legal texts, where high-quality corpora and relevance judgments are scarce. To address this limitation, we propose STCALIR, a framework for generating semi-synthetic test collections directly from raw legal documents. The pipeline follows the Cranfield paradigm, maintaining its core components of topics, corpus, and relevance judgments, while significantly reducing manual effort through automated multi-stage retrieval and filtering, achieving a 99% reduction in annotation workload. We validate STCALIR using the Mr. TyDi benchmark, demonstrating that the resulting semi-synthetic relevance judgments yield retrieval effectiveness comparable to human-annotated evaluations (Hit@10 \approx 0.785). Furthermore, system-level rankings derived from these labels exhibit strong concordance with human-based evaluations, as measured by Kendall's τ (0.89) and Spearman's \r{ho} (0.92). Overall, STCALIR offers a reproducible and cost-efficient solution for constructing reliable test collections in low-resource legal domains.

法律信息检索半合成数据低资源场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。