arXiv:2607.05582cs.IR2026-07中稿 · the ASAIL Workshop…

用提示词比微调更高效地排序判例句,解释法律条文。

Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval

论文配图:Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval
图 1 · 摘自论文原文
  • 用解码器模型零样本提示,不需训练直接推理。
  • 最佳系统在42个法律概念上超越已有最优结果。
  • 适合法律人工智能、法律信息检索研究者参考。

法律条文中常使用模糊术语,实务中常依赖判例来解释。本文研究通过判例句子的有用性对法律概念或目标条文进行排序的任务,使用包含42个美国法典概念、26,959条句子的标注数据集,分为四类解释价值。比较两类方法:(i) 编码器模型(ModernBERT)的监督微调;(ii) 解码器模型的零样本提示。结果显示,ModernBERT在所有概念和标准NDCG截断值下与早期BERT系列基线表现相当。相比之下,提示解码器模型整体效果最优,最佳系统在该任务上超越所有先前报告的最先进结果。

原文摘要 · Abstract (English)

Legal concepts in statutes are often expressed using vague terms, and practitioners frequently turn to case law to interpret them. We study the task of ranking case-law sentences by their usefulness for explaining a concept or target statutory term, using an established dataset of 26,959 sentences covering 42 U.S. Code concepts labeled into four explanatory-value categories. We compare two families of methods: (i) supervised fine-tuning of encoder-only models (ModernBERT) and (ii) zero-shot prompting of decoder-only models. We show that across all concepts and standard NDCG cutoffs, ModernBERT largely matches earlier BERT-family baselines. In contrast, prompting decoder-only models achieves the strongest overall effectiveness, with our best system surpassing all previously reported state-of-the-art results on this task.

法律AI提示工程信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。