探究法律条文解释检索中人工标注的必要性与替代方案
Are manual annotations necessary for statutory interpretations retrieval?
- 实验测试不同标注数量对模型性能的影响
- 对比随机选句与精选候选句标注的效果差异
- 评估大模型自动标注的可行性与效果
法律研究中,识别法官对法律概念的解释至关重要,这些解释可作为判例参考或帮助公众理解法律。当前主流方法依赖句子排序和基于人工标注样本训练的语言模型,但该过程成本高且需为每个概念重复进行。本文通过多项实验探究人工标注的必要性:首先测试每类法律概念最优标注数量;其次比较随机选句与仅标注高质量候选句的效果差异;最后评估利用大语言模型(LLM)自动化标注的成效。
原文摘要 · Abstract (English)
One of the elements of legal research is looking for cases where judges have extended the meaning of a legal concept by providing interpretations of what a concept means or does not mean. This allow legal professionals to use such interpretations as precedents as well as laymen to better understand the legal concept. The state-of-the-art approach for retrieving the most relevant interpretations for these concepts currently depends on the ranking of sentences and the training of language models over annotated examples. That manual annotation process can be quite expensive and need to be repeated for each such concept, which prompted recent research in trying to automate this process. In this paper, we highlight the results of various experiments conducted to determine the volume, scope and even the need for manual annotation. First of all, we check what is the optimal number of annotations per a legal concept. Second, we check if we can draw the sentences for annotation randomly or there is a gain in the performance of the model, when only the best candidates are annotated. As the last question we check what is the outcome of automating the annotation process with the help of an LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。