大模型会无意识泄露地理位置,连用‘未知’占位符也难逃影响。
Unintended Effects of Geographic Conditioning in Large Language Models

- 用用户位置信息做条件输入,模型却产生区域偏见。
- 泄漏率最高飙升793倍,如Llama 3.1-8B从0.04%升至31.7%。
- 连‘未知’占位符都引发泄漏,说明身份标签本身就有诱导性。
现代对话式AI系统常依赖用户元数据来定位回复,但由此引入的隐性区域偏差仍不清楚。本文评估了位置泄漏现象:即使用户提示地理中立,模型仍会生成地理相关表述。在创意写作和开放问答任务中,即使是顶尖大模型,在暴露于位置元数据时也系统性偏好特定地区输出,泄漏率最高提升793倍(如Llama 3.1-8B从0.04%增至31.7%,Qwen3-8B为21.3%,Claude Sonnet 4.6为8.8%)。分析发现一种新型结构化条件效应:将注入位置替换为占位符“Unknown”后,泄漏率仍比基线高72倍,表明用户资料框架本身——无论是否含具体地理内容——已构成生成引导信号。
原文摘要 · Abstract (English)
Modern conversational AI systems frequently rely on user metadata to localize responses, yet the unintended regional biases introduced by this hidden context remain poorly understood. In this work, we evaluate location leakage: the phenomenon where a model generates geographic references despite receiving a geographically neutral user prompt. Across both creative writing and open-ended Q&A prompts, even state-of-the-art LLMs systematically favor region-specific outputs when exposed to location metadata, with leakage spiking by up to 793 times above baseline (e.g., from 0.04% to 31.7% for Llama 3.1-8B, and 21.3% and 8.8% for Qwen3-8B and Claude Sonnet 4.6, respectively). Our analysis further shows a novel structural conditioning effect: replacing the injected location with the placeholder "Unknown" still elevates leakage by up to 72 times above baseline, demonstrating that the user profile frame itself, independent of any geographic content, acts as a generative conditioning signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。