arXiv:2512.13654cs.CLcs.AI2025-12

研究大模型在判例分类中的记忆能力,发现带提示的模型更稳定准确。

Large-Language Memorization During the Classification of United States Supreme Court Cases

  • 用提示工程增强模型记忆,提升判例分类性能
  • 在15类和279类任务上均比BERT模型高约2分
  • 适合关注法律AI、模型记忆机制的研究者

大型语言模型(LLMs)在问答之外的任务中表现出多样化响应,有时被称作“幻觉”。为理解模型如何回应,研究其记忆策略至关重要。本文深入分析基于美国最高法院(SCOTUS)判决的分类任务,该数据集因句子过长、法律术语复杂、结构非标准及领域词汇密集而极具挑战性。实验采用最新的微调与检索方法,包括参数高效微调、自动建模等,在两个传统类别分类任务(15类与279类)上进行测试。结果表明,基于提示的记忆型模型(如DeepSeek)在两项任务上均优于此前的BERT基线模型,表现提升约2分,展现出更强的鲁棒性。

原文摘要 · Abstract (English)

Large-language models (LLMs) have been shown to respond in a variety of ways for classification tasks outside of question-answering. LLM responses are sometimes called "hallucinations" since the output is not what is ex pected. Memorization strategies in LLMs are being studied in detail, with the goal of understanding how LLMs respond. We perform a deep dive into a classification task based on United States Supreme Court (SCOTUS) decisions. The SCOTUS corpus is an ideal classification task to study for LLM memory accuracy because it presents significant challenges due to extensive sentence length, complex legal terminology, non-standard structure, and domain-specific vocabulary. Experimentation is performed with the latest LLM fine tuning and retrieval-based approaches, such as parameter-efficient fine-tuning, auto-modeling, and others, on two traditional category-based SCOTUS classification tasks: one with 15 labeled topics and another with 279. We show that prompt-based models with memories, such as DeepSeek, can be more robust than previous BERT-based models on both tasks scoring about 2 points better than previous models not based on prompting.

大模型法律AI分类任务记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。