测试大模型在沙特与欧盟数据法上的引用造假问题,发现对沙特法造假率高达67%。
Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR
- 构建中英双语120题基准,检验模型对欧盟与沙特数据法的引用准确性。
- 沙特PDPL引用错误率达60%-77%,主要因法规与条例混淆,91%假引信心超0.8。
- 模型误判与语言无关,高置信度不能防错,需人工核验才能用于合规审查。
组织与监管机构日益依赖大语言模型(LLMs)解答合规问题,但错误的法律条文引用可能悄然传播至法律建议、合规文件和政策决策中。本文提出一个包含120个问题的双语基准,评估免费可访问的LLMs在欧盟《通用数据保护条例》(GDPR)与沙特《个人数据保护法》(PDPL)下的引用能力。测试涵盖直接引用检索、虚假前提验证及故意无解的“陷阱”问题(如已废止条款或仅存在于实施条例中的截止日期)。所有问题均以阿拉伯语和英语提出,评分基于人工校验的黄金标准。评估三款模型(Gemini 2.5 Flash、GPT-OSS-120B、Nemotron-3-Super-120B)发现显著司法差异:对GDPR的直接引用准确率达94%-100%,而对沙特PDPL的引用错误率为60%-77%,且不受查询语言影响;最高伪造率(67%)源于法规与条例混淆,91%的错误引用被模型以>=0.8的信心声称。引用错误与法律管辖地相关,而非查询语言,模型自信无法提供防护,表明机构在使用大模型进行合规筛查时,必须依赖文本比对验证,而非模型自我判断。
原文摘要 · Abstract (English)
Organizations and regulators increasingly consult large language models (LLMs) for regulatory-compliance questions, yet a wrong statutory citation can silently propagate into legal advice, compliance documentation, and policy decisions. We introduce a bilingual benchmark of 120 questions probing whether freely accessible LLMs fabricate article citations for two data-protection instruments: the EU General Data Protection Regulation (GDPR) and the Saudi Personal Data Protection Law (PDPL). The benchmark pairs direct citation retrieval questions with false premise verification probes and deliberately unanswerable "trap" questions -- including questions about a repealed article and about deadlines that exist only in implementing regulations, not in the law itself. Every question is posed in both Arabic and English, and all scoring is fully automatic against a manually verified gold reference. Evaluating three freely accessible models (Gemini 2.5 Flash, GPT-OSS-120B, Nemotron-3-Super-120B), we find a dramatic jurisdiction gap: near-ceiling citation accuracy on the GDPR (94-100% on direct retrieval) against majority fabrication on the Saudi PDPL (60-77%), invariant to query language; the highest fabrication rates (67%) arise from statute-vs-regulations confusion, and 91% of fabricated citations are asserted with confidence >= 0.8. Fabrication tracks the jurisdiction of the law, not the language of the query, and model confidence provides no protection -- indicating that verbatim-verification safeguards, rather than model self confidence, must gate any institutional reliance on LLMs for compliance screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。