用大模型辅助审查德语劳动合同条款合法性,提升法律文本理解效率。
LLMs for Legal Subsumption in German Employment Contracts
- 利用大模型结合上下文学习,判断合同条款是否合法、不公平或无效。
- 使用法规摘要后,对无效条款的召回率和加权F1达80%,显著提升。
- 适合法律科技从业者及需自动化合同审查的研究者参考。
法律工作以文本密集和资源消耗大为特点,为自然语言处理研究带来独特挑战与机遇。尽管数据驱动方法已取得进展,但其可解释性与可信度不足,限制了在动态法律环境中的应用。为此,我们与法律专家合作,扩展了现有数据集,并探索使用大语言模型(LLMs)及上下文学习,评估德语雇佣合同中条款的合法性。研究对比了三种法律情境下的表现:无法律上下文、完整法律法规与判例全文、以及提炼后的法规摘要(称为考试指南)。结果表明,完整文本适度提升性能,而考试指南显著提高无效条款的召回率和加权F1分数,最高达80%。尽管如此,使用完整文本时,大模型的表现仍远低于人类律师。我们贡献了一个扩展数据集,包含考试指南、引用法律来源及对应标注,并公开代码与所有日志文件。研究揭示了大模型在协助律师审查合同合法性方面的潜力,同时也指出了当前方法的局限性。
原文摘要 · Abstract (English)
Legal work, characterized by its text-heavy and resource-intensive nature, presents unique challenges and opportunities for NLP research. While data-driven approaches have advanced the field, their lack of interpretability and trustworthiness limits their applicability in dynamic legal environments. To address these issues, we collaborated with legal experts to extend an existing dataset and explored the use of Large Language Models (LLMs) and in-context learning to evaluate the legality of clauses in German employment contracts. Our work evaluates the ability of different LLMs to classify clauses as "valid," "unfair," or "void" under three legal context variants: no legal context, full-text sources of laws and court rulings, and distilled versions of these (referred to as examination guidelines). Results show that full-text sources moderately improve performance, while examination guidelines significantly enhance recall for void clauses and weighted F1-Score, reaching 80\%. Despite these advancements, LLMs' performance when using full-text sources remains substantially below that of human lawyers. We contribute an extended dataset, including examination guidelines, referenced legal sources, and corresponding annotations, alongside our code and all log files. Our findings highlight the potential of LLMs to assist lawyers in contract legality review while also underscoring the limitations of the methods presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。