构建法律事实核查基准,验证普通人说法是否符合最高法院判例。
CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval
- 用大模型从专家案例摘要生成普通人说法,跨越法律术语与日常语言的鸿沟。
- 数据集含6294个说法,分支持、反驳、推翻三类,考虑判例时效性。
- 发现联网搜索反而降低准确率,因引入噪声和非权威判例,适合法律AI研究者。
自动化事实核查主要针对通用知识与静态语料库,忽视了法律等高风险领域中动态且复杂的真理。我们提出CaseFacts,一个用于验证美国最高法院判例中民间法律主张的基准数据集。不同于现有资源在正式文本间映射,CaseFacts要求系统弥合普通人陈述与专业法律术语之间的语义差距,并考虑判例的时间有效性。该数据集包含6,294条主张,分为支持、反驳或推翻三类。通过多阶段流程构建:利用大语言模型(LLMs)从专家案例摘要生成主张,并采用新颖的语义相似性启发式方法高效识别并验证复杂法律推翻关系。对先进大模型的实验表明,该任务仍具挑战性;值得注意的是,使用无限制网络搜索的模型性能低于封闭书本基线,因检索到大量噪声且非权威判例。我们发布CaseFacts以推动法律事实核查系统的研究。
原文摘要 · Abstract (English)
Automated Fact-Checking has largely focused on verifying general knowledge against static corpora, overlooking high-stakes domains like law where truth is evolving and technically complex. We introduce CaseFacts, a benchmark for verifying colloquial legal claims against U.S. Supreme Court precedents. Unlike existing resources that map formal texts to formal texts, CaseFacts challenges systems to bridge the semantic gap between layperson assertions and technical jurisprudence while accounting for temporal validity. The dataset consists of 6,294 claims categorized as Supported, Refuted, or Overruled. We construct this benchmark using a multi-stage pipeline that leverages Large Language Models (LLMs) to synthesize claims from expert case summaries, employing a novel semantic similarity heuristic to efficiently identify and verify complex legal overrulings. Experiments with state-of-the-art LLMs reveal that the task remains challenging; notably, augmenting models with unrestricted web search degrades performance compared to closed-book baselines due to the retrieval of noisy, non-authoritative precedents. We release CaseFacts to spur research into legal fact verification systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。