大模型常误判人名,隐私保护存在语言陷阱
Can Large Language Models Really Recognize Your Name?
- 构建12000+模糊人名基准,测试模型识别能力
- 真实人名识别率下降20%~40%,特定场景下忽略率升至4倍
- 适合关注隐私安全与模型公平性的研究者阅读
大型语言模型(LLM)被广泛用于隐私保护中检测敏感信息泄露,其前提假设是模型能可靠识别人名。本文揭示,由于上下文语言线索模糊,即使在短文本中,大模型也频繁误判各类人名。我们基于姓名规律性偏差现象构建了AmBench基准,包含超过12000个真实但具有歧义的人名,每条人名出现在数十个语义多义的简短文本片段中。对12个先进大模型的实验显示,与易识别人名相比,AmBench人名的召回率下降20%至40%。当上下文包含良性提示注入(即类似指令的用户文本)时,在Anthropic AI用于从用户对话中提取隐私保护洞察的Clio工具中,这些模糊人名被忽略的可能性增加四倍。研究揭示了基于大模型的隐私解决方案在性能与公平性上的盲点,呼吁系统性探究其隐私失效模式与应对策略。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used in privacy pipelines to detect and remedy sensitive data leakage. These solutions often rely on the premise that LLMs can reliably recognize human names, one of the most important categories of personally identifiable information (PII). In this paper, we reveal how LLMs can consistently mishandle broad classes of human names even in short text snippets due to ambiguous linguistic cues in the contexts. We construct AmBench, a benchmark of over 12,000 real yet ambiguous human names based on the name regularity bias phenomenon. Each name appears in dozens of concise text snippets that are compatible with multiple entity types. Our experiments with 12 state-of-the-art LLMs show that the recall of AmBench names drops by 20--40% compared to more recognizable names. This uneven privacy protection due to linguistic properties raises important concerns about the fairness of privacy enforcement. When the contexts contain benign prompt injections -- instruction-like user texts that can cause LLMs to conflate data with commands -- AmBench names can become four times more likely to be ignored in Clio, an LLM-powered enterprise tool used by Anthropic AI to extract supposedly privacy-preserving insights from user conversations with Claude. Our findings showcase blind spots in the performance and fairness of LLM-based privacy solutions and call for a systematic investigation into their privacy failure modes and countermeasures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。