arXiv:2512.05681cs.CLcs.AI2025-12中稿 · presentation as a …

对比两种嵌入模型在嘈杂司法标签下的判例检索效果

Retrieving Semantically Similar Decisions under Noisy Institutional Labels: Robust Comparison of Embedding Methods

  • 用滑动窗口和注意力池化训练领域专用BERT,对比通用OpenAI嵌入
  • OpenAI模型在@10/@20/@100指标上显著优于领域BERT,差异具统计意义
  • 框架可应对标签噪声,适合历史司法数据的评估场景

判例检索主要依赖数据库查询,耗时且效率低。本文针对捷克宪法法院判决,在三种不同设置下对比两种模型:(i) 通用大规模嵌入模型(OpenAI),(ii) 基于约3万份判决文本、从零开始训练的领域专用BERT,采用滑动窗口与注意力池化。提出一种抗噪声评估方法,包括基于IDF加权关键词重叠的分级相关性判断,通过两个阈值(0.20平衡、0.28严格)进行二值化,使用配对自助法检验显著性,并结合nDCG诊断与定性分析。尽管绝对nDCG表现平平(预期因标签噪声所致),通用OpenAI嵌入在两种阈值下于@10/@20/@100指标均显著优于领域预训练BERT,差异具有统计显著性。诊断表明低绝对性能源于标签漂移与强理想主义,而非模型无用。此外,该框架具备鲁棒性,适用于存在噪声真值数据的评估,这在来自异构司法数据库的复杂数据中十分常见。

原文摘要 · Abstract (English)

Retrieving case law is a time-consuming task predominantly carried out by querying databases. We provide a comparison of two models in three different settings for Czech Constitutional Court decisions: (i) a large general-purpose embedder (OpenAI), (ii) a domain-specific BERT-trained from scratch on ~30,000 decisions using sliding windows and attention pooling. We propose a noise-aware evaluation including IDF-weighted keyword overlap as graded relevance, binarization via two thresholds (0.20 balanced, 0.28 strict), significance via paired bootstrap, and an nDCG diagnosis supported with qualitative analysis. Despite modest absolute nDCG (expected under noisy labels), the general OpenAI embedder decisively outperforms the domain pre-trained BERT in both settings at @10/@20/@100 across both thresholds; differences are statistically significant. Diagnostics attribute low absolutes to label drift and strong ideals rather than lack of utility. Additionally, our framework is robust enough to be used for evaluation under a noisy gold dataset, which is typical when handling data with heterogeneous labels stemming from legacy judicial databases.

判例检索嵌入模型噪声标签司法AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。