构建法律事件时间排序基准,评估大模型在真实法律文本中的时序理解能力。
LexTime: A Benchmark for Temporal Ordering of Legal Events
- 基于美国联邦诉状构建512个事件对,标注时间关系用于评测。
- 模型在法律文本上准确率比叙事文本高10.5%,隐含事件对达80.8%。
- 法律语言复杂性仍是难点,适合法律AI研究者参考。
理解时间关系并准确重建事件时间线对判例分析、合规监控和法律摘要至关重要。然而现有基准缺乏针对法律语言的评估,导致大模型在法律时序推理方面的能力评估存在空白。我们提出LexTime,一个专为评估大模型法律事件排序能力设计的数据集,包含512个来自美国联邦诉状的实例,每个实例标注了事件对及其时间关系。研究发现:(1)大模型在法律事件排序上的表现优于叙事文本,最高提升10.5%;(2)更长输入上下文和隐含事件可提高准确率,隐含-显式事件对达到80.8%;(3)法律语言复杂性和嵌套从句仍是主要挑战。尽管性能可观,法律文本特有特征仍是时序推理瓶颈,我们提出了具体的建模改进方向。
原文摘要 · Abstract (English)
Understanding temporal relationships and accurately reconstructing the event timeline is important for case law analysis, compliance monitoring, and legal summarization. However, existing benchmarks lack specialized language evaluation, leaving a gap in understanding how LLMs handle event ordering in legal contexts. We introduce LexTime, a dataset designed to evaluate LLMs' event ordering capabilities in legal language, consisting of 512 instances from U.S. Federal Complaints with annotated event pairs and their temporal relations. Our findings show that (1) LLMs are more accurate on legal event ordering than on narrative texts (up to +10.5%); (2) longer input contexts and implicit events boost accuracy, reaching 80.8% for implicit-explicit event pairs; (3) legal linguistic complexities and nested clauses remain a challenge. While performance is promising, specific features of legal texts remain a bottleneck for legal temporal event reasoning, and we propose concrete modeling directions to better address them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。