测试大模型在法律知识注入攻击下的鲁棒性,发现其易受错字误导。
J&H: Evaluating the Robustness of Large Language Models Under Knowledge-Injection Attacks in Legal Domain
- 设计法律知识注入攻击,逐层干扰推理前提与结论生成
- 多数大模型因错字等细微错误产生错误判断,逻辑推理能力弱
- 适合关注法律AI安全性的研究者与从业者参考
随着大语言模型(LLMs)规模和能力的提升,其在法律等知识密集型领域的应用日益受到关注。然而,这些模型是否真正基于领域知识进行推理仍存疑。若模型仅依赖特定词汇或模式而非语言底层逻辑做判断,将带来‘大模型作为裁判’在实际应用中的重大风险。为此,本文提出一种法律知识注入攻击方法,用于评估大模型在法律领域的鲁棒性,以推断其是否掌握法律知识与推理逻辑。我们构建了J&H框架,针对法律任务中的主要前提、次要前提及结论生成环节实施攻击。收集了真实司法实践中专家可能犯的错误,如拼写错误、法律同义词使用、外部法条检索不准等。现实中法律专家通常忽略此类错误,依赖逻辑判断;但大模型却易被误导。我们在通用及领域专用大模型上实施攻击,结果显示当前模型对本实验中的攻击均不鲁棒。此外,我们提出并对比了几种提升知识鲁棒性的方法。
原文摘要 · Abstract (English)
As the scale and capabilities of Large Language Models (LLMs) increase, their applications in knowledge-intensive fields such as legal domain have garnered widespread attention. However, it remains doubtful whether these LLMs make judgments based on domain knowledge for reasoning. If LLMs base their judgments solely on specific words or patterns, rather than on the underlying logic of the language, the ''LLM-as-judges'' paradigm poses substantial risks in the real-world applications. To address this question, we propose a method of legal knowledge injection attacks for robustness testing, thereby inferring whether LLMs have learned legal knowledge and reasoning logic. In this paper, we propose J&H: an evaluation framework for detecting the robustness of LLMs under knowledge injection attacks in the legal domain. The aim of the framework is to explore whether LLMs perform deductive reasoning when accomplishing legal tasks. To further this aim, we have attacked each part of the reasoning logic underlying these tasks (major premise, minor premise, and conclusion generation). We have collected mistakes that legal experts might make in judicial decisions in the real world, such as typos, legal synonyms, inaccurate external legal statutes retrieval. However, in real legal practice, legal experts tend to overlook these mistakes and make judgments based on logic. However, when faced with these errors, LLMs are likely to be misled by typographical errors and may not utilize logic in their judgments. We conducted knowledge injection attacks on existing general and domain-specific LLMs. Current LLMs are not robust against the attacks employed in our experiments. In addition we propose and compare several methods to enhance the knowledge robustness of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。