arXiv:2511.06700cs.CYcs.AI2025-11被引 2

比较不同地区法律大模型幻觉率,发现其表现随地理位置显著变化。

Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries

  • 基于功能主义构建跨地域法律对比方法
  • 洛杉矶、伦敦、悉尼三地幻觉率差异显著
  • 适合关注法律AI地理偏差的研究者与使用者

如何衡量大语言模型在不同地区法律知识上的差异?本研究提出基于功能主义的比较方法,构建来自Reddit用户求助帖的法律事实场景数据集,涵盖家庭、住房、就业、犯罪和交通问题。针对洛杉矶、伦敦和悉尼三地,让主流闭源大模型生成相关法律条文摘要,并人工评估幻觉率。结果显示,主流大模型的法律信息幻觉率显著依赖于地理位置,表明其法律服务能力存在明显地域不均。此外,幻觉率与模型多次采样中多数响应频率呈强负相关,反映模型对法律事实预测的不确定性。

原文摘要 · Abstract (English)

How do we make a meaningful comparison of a large language model's knowledge of the law in one place compared to another? Quantifying these differences is critical to understanding if the quality of the legal information obtained by users of LLM-based chatbots varies depending on their location. However, obtaining meaningful comparative metrics is challenging because legal institutions in different places are not themselves easily comparable. In this work we propose a methodology to obtain place-to-place metrics based on the comparative law concept of functionalism. We construct a dataset of factual scenarios drawn from Reddit posts by users seeking legal advice for family, housing, employment, crime and traffic issues. We use these to elicit a summary of a law from the LLM relevant to each scenario in Los Angeles, London and Sydney. These summaries, typically of a legislative provision, are manually evaluated for hallucinations. We show that the rate of hallucination of legal information by leading closed-source LLMs is significantly associated with place. This suggests that the quality of legal solutions provided by these models is not evenly distributed across geography. Additionally, we show a strong negative correlation between hallucination rate and the frequency of the majority response when the LLM is sampled multiple times, suggesting a measure of uncertainty of model predictions of legal facts.

法律AI幻觉率地理偏差大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。