arXiv:2510.01792cs.CLcs.AI2025-10被引 1

对比16种无监督指标,评估从俄语判决书提取法律信息的效果。

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

  • 用16种无监督方法评估7类法律信息块的提取质量,无需人工标注。
  • 词频一致性(TF Coherence)与覆盖率在专家评分上相关性最高(皮尔逊r=0.54)。
  • 大模型评分虽可测但相关性中等,不能替代人工判断,适合快速筛查。

人工智能在法律自然语言处理中的快速发展,亟需可扩展的文本提取评估方法。本研究评估了16种无监督指标,包括新提出的方法,用于衡量从1,000份匿名俄语司法判决书中提取7个语义块的质量,该评估基于7,168条专家评分(1–5分制)。这些指标涵盖文档级、语义、结构、伪真值及法律专用类别,均无需预标注真值。通过自助法相关性、林氏一致性相关系数(CCC)和平均绝对误差(MAE)分析发现,词频一致性(Pearson $r = 0.540$,Lin CCC = 0.512,MAE = 0.127)和覆盖率/块完整性(Pearson $r = 0.513$,Lin CCC = 0.443,MAE = 0.139)与专家评分最一致;而法律术语密度则呈现强负相关(Pearson $r = -0.479$,Lin CCC = -0.079,MAE = 0.394)。LLM评估得分(均值=0.849,Pearson $r = 0.382$,Lin CCC = 0.325,MAE = 0.197)表现中等,使用gpt-4.1-mini通过g4f实现,表明其对法律文本适配有限。结果表明,尽管无监督指标支持可扩展筛选,但因相关性中等、一致性低,仍无法完全替代高风险法律场景下的人工判断。本工作为法律NLP提供了无需标注的评估工具,对司法分析与伦理化AI部署具有意义。

原文摘要 · Abstract (English)

The rapid advancement of artificial intelligence in legal natural language processing demands scalable methods for evaluating text extraction from judicial decisions. This study evaluates 16 unsupervised metrics, including novel formulations, to assess the quality of extracting seven semantic blocks from 1,000 anonymized Russian judicial decisions, validated against 7,168 expert reviews on a 1--5 Likert scale. These metrics, spanning document-based, semantic, structural, pseudo-ground truth, and legal-specific categories, operate without pre-annotated ground truth. Bootstrapped correlations, Lin's concordance correlation coefficient (CCC), and mean absolute error (MAE) reveal that Term Frequency Coherence (Pearson $r = 0.540$, Lin CCC = 0.512, MAE = 0.127) and Coverage Ratio/Block Completeness (Pearson $r = 0.513$, Lin CCC = 0.443, MAE = 0.139) best align with expert ratings, while Legal Term Density (Pearson $r = -0.479$, Lin CCC = -0.079, MAE = 0.394) show strong negative correlations. The LLM Evaluation Score (mean = 0.849, Pearson $r = 0.382$, Lin CCC = 0.325, MAE = 0.197) showed moderate alignment, but its performance, using gpt-4.1-mini via g4f, suggests limited specialization for legal textse. These findings highlight that unsupervised metrics, including LLM-based approaches, enable scalable screening but, with moderate correlations and low CCC values, cannot fully replace human judgment in high-stakes legal contexts. This work advances legal NLP by providing annotation-free evaluation tools, with implications for judicial analytics and ethical AI deployment.

法律NLP无监督评估文本提取大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。