多语言大模型会默认用输入语言判断法律适用地区,可能误导用户。
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs

- 以输入语言作为默认法律管辖地判断依据
- 中文输入时53.3%给出中国法规,英文输入时74.5%倾向美国法规
- 建议接口明确询问或声明回答所基于的司法辖区
大语言模型越来越多地回答税收、劳工保护、医疗、教育、养老金及行政程序等问题,其有效性常取决于适用的司法辖区。多语言用户可能使用自己最熟悉的语言提问,而非与实际管辖地对应的语言。我们研究部署中的大模型在未指定国家或地区时,是否将输入语言作为默认的司法辖区信号。评估了七款在美国或中国开发的模型,在英语和中文下对60个未指定司法辖区的法律行政问题进行测试,共获得2,520条人工标注响应。结果显示,中文输入更常生成中国特定答案,英文输入则更多采用美国、比较性或通用答案。要求单一答案的提示进一步强化此趋势:综合所有模型,74.5%的英文输入响应采用美国框架,53.3%的中文输入响应采用中国框架。该方向性模式在所有七款模型中均存在。我们将其称为制度框架误选风险——流利的回答可能依赖于用户未意图的法律行政背景,尤其当用户偏好语言与目标司法辖区不一致时。大模型接口不应仅根据输入语言分配制度建议;当位置缺失时,应主动询问或声明答案的司法范围。
原文摘要 · Abstract (English)
LLMs increasingly answer questions about taxes, labor protections, healthcare, education, pensions, and administrative procedures, where usefulness often depends on the applicable jurisdiction. Multilingual users may write in their most comfortable language rather than one associated with the country or region whose rules apply. We ask whether deployed LLMs use input language as a default jurisdictional signal when prompts omit any country or region. Prior multilingual audits show that prompt language can shift cultural, political, or normative outputs; we examine which legal-administrative framework models supply when jurisdiction is underspecified. We evaluate seven LLMs developed in the United States or China on 60 underspecified legal-administrative prompts in English and Mandarin Chinese under three system-prompt conditions, yielding 2,520 manually annotated responses. Across models and conditions, Chinese input more often produces China-specific answers, while English input more often produces U.S.-specific, comparative, or generic answers. Prompts requiring a single answer further increase jurisdiction selection: pooled across models, 74.5% of English-input responses adopt a U.S. framework, while 53.3% of Chinese-input responses adopt a China framework. This directional pattern appears in all seven models. We describe this deployment-level pattern as institutional-framework misselection risk: a fluent answer may rely on a legal-administrative context the user did not intend, especially when their preferred language differs from the relevant jurisdiction. LLM interfaces should not route institutional advice by input language alone; when location is absent, they should request it or state the jurisdictional scope of the answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。