解决大模型在法律问答中因时间失效导致的过时与偏新问题
Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

- 构建312个带时间标签的德国法律问答数据集,覆盖三类时效性问题
- 检索增强生成显著提升准确率,而直接搜索易产生新法偏好偏差
- 强调法律问答必须将时间有效性作为硬约束,适合法律AI研发者参考
大语言模型在法律研究中应用日益广泛,但其固定训练截止时间和静态参数知识与法律动态演变不匹配。本文研究两类时间失效模式:截止后过时(模型使用已修改的旧法规)和近期偏差(即使历史版本适用,仍倾向选择新条款)。为此,我们构建了312个专家验证的时间敏感德国法律问答对,涵盖三类:截止后修订问题、修订前问题、多条款修订前问题。评估OpenAI、Anthropic和DeepSeek的五款LLM在四种推理设置下表现:原生模式、网络搜索,以及两种通过事实日期提取和版本过滤实现时间有效性约束的检索增强变体。采用基于大模型的评判标准,经人类专家评分验证,发现原生模式在截止后严重退化。两种RAG方法在所有题型上均有显著提升,而网络搜索结果不稳定,且在历史锚定任务中表现出明显近期偏好。结果表明,可靠法律问答必须将时间有效性视为硬约束。
原文摘要 · Abstract (English)
Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, where models apply superseded rules after legislative amendments, and recency bias, where models prefer newer provisions even when a historical version governs the fact pattern. To this end, we present a benchmark of 312 expert-validated, time-sensitive German statutory QA pairs spanning three categories: Post-Cutoff Amendment Questions, Pre-Amendment Questions, and Multi-Provision Pre-Amendment Questions. We evaluate five LLMs by OpenAI, Anthropic and DeepSeek under four inference settings: Vanilla, Web-search, and two retrieval-augmented variants that enforce temporal validity via a fact date extraction and version filtering. Using an LLM-as-a-judge validated against human expert ratings, we find severe degradation in the Vanilla post-cutoff setting. Both RAG approaches substantially improve performance across all question types, while web search yields unstable gains and exhibits a marked recency bias on historically anchored tasks. Our results indicate that reliable legal QA requires treating temporal validity as a hard constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。