AI在法律引用格式上表现不佳,但结合规则引擎可显著提升准确率。
Bye-bye, Bluebook? Automating Legal Drudgery With AI-Augmented Rule Following
- 用神经符号系统先解析文本再执行规则,提升格式准确性。
- 顶级模型零样本下仅42.6%合规,远低于人类编辑水平。
- 纯检索增强效果有限,规则执行需显式编程支持。
法律AI承诺自动化律师的重复性工作,但其实际表现尚不明确。本文首次实证检验了前沿语言模型在最常见法律文书任务——蓝皮书(Bluebook)引用格式上的表现。构建了包含2,058个查询的新基准,发现主流模型在零样本条件下仅42.6%生成完全合规引用。对五家顶尖法学期刊的实验显示,即使使用“推理”模型,其表现仍远低于期刊编辑选拔赛中人类候选人的平均分。提供规则虽略有改善,但检索增强生成(RAG)无法确保规则遵循。本文提出一种神经符号系统:先由模型提取结构化引用元素,再交由确定性规则引擎格式化,使平均准确率提升32.4个百分点,最高达85.5%。结果表明,仅靠大模型无法实现法律自动化,但结合符号规则引擎可提供可行路径。
原文摘要 · Abstract (English)
One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without calling for much discretion. Yet it remains an open question how well AI models actually perform on such tasks. This article presents the first empirical examination of AI performance on perhaps the most ubiquitous and lamented form of legal drudgery: citation formatting under the Bluebook. We make four contributions. First, we develop a new benchmark of 2,058 Bluebook queries and show that, on average, frontier language models produce a fully compliant legal citation only 42.6% of the time in a zero-shot setting. Second, we conduct an experiment with five top law reviews and show that even a "reasoning" model falls far below the average score of the human candidates in these journals' annual editor-selection competitions. Third, we show that simply providing the models with the rules offers only modest improvements, calling into question the ability of retrieval-augmented generation (RAG) to ensure rule-following alone. Finally, we develop an approach that does meaningfully improve compliance: a neuro-symbolic system that first uses a model to parse natural language into structured citation elements, and then delegates the formatting to a deterministic rule-execution engine. This approach achieves an average accuracy increase of 32.4 percentage points and total accuracy of up to 85.5% on our benchmark. These results point toward a reorientation for legal AI. The original promise of automating drudgery still remains out of reach for even frontier language models on their own -- but pairing them with symbolic rule engines may offer a tractable path forward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。