arXiv:2412.15902cs.CLcs.AI2024-12被引 3

评测开源大模型在德语法律教育中的分析能力,发现其在简单任务中表现良好,复杂任务则需增强提示策略。

On the Suitability of pre-trained foundational LLMs for Analysis in German Legal Education

  • 采用检索增强生成的提示示例选择方法提升高数据场景下的预测效果
  • 在论点挖掘与作文评分任务中表现优于基础模型,零样本下仍具优势
  • 复杂法律意见分析仍难胜任,需结合外部知识与精心设计提示

我们证明当前开源的基础大模型具备足够的指令理解能力与德语法律背景知识,可在教育场景中完成部分法律分析任务。然而,在特定任务如「Gutachtenstil」评估风格分类,或处理完整法律意见等复杂上下文时,模型性能显著下降。即使使用扩展上下文和有效提示策略,也难以超越词袋(Bag-of-Words)基线。为此,我们提出一种基于检索增强生成的提示示例选择方法,在高数据可用性场景下显著提升预测表现。进一步评估显示,预训练大模型在论点挖掘与自动作文评分两个标准任务上表现更佳。整体而言,在标注数据极少或无标注数据情况下,预训练大模型结合思维链(Chain-of-Thought)提示可明显优于基线。

原文摘要 · Abstract (English)

We show that current open-source foundational LLMs possess instruction capability and German legal background knowledge that is sufficient for some legal analysis in an educational context. However, model capability breaks down in very specific tasks, such as the classification of "Gutachtenstil" appraisal style components, or with complex contexts, such as complete legal opinions. Even with extended context and effective prompting strategies, they cannot match the Bag-of-Words baseline. To combat this, we introduce a Retrieval Augmented Generation based prompt example selection method that substantially improves predictions in high data availability scenarios. We further evaluate the performance of pre-trained LLMs on two standard tasks for argument mining and automated essay scoring and find it to be more adequate. Throughout, pre-trained LLMs improve upon the baseline in scenarios with little or no labeled data with Chain-of-Thought prompting further helping in the zero-shot case.

法律AI大模型评估提示工程德语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。