arXiv:2605.12147cs.CRcs.LG2026-05

用1000人真实数据测试大模型能否模拟个人隐私行为。

PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior

论文配图:PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior
图 1 · 摘自论文原文
  • 基于用户人口统计、经历和态度三类特征,测试大模型隐私决策模拟能力。
  • 最强模型仅40.4%准确率,远未达到真实水平。
  • 有经验但态度不谨慎的用户最难模拟,适合安全与隐私研究者参考。

大型语言模型(LLMs)在模拟人类行为方面应用日益广泛,但其对个体隐私决策的模拟能力尚不明确。本文提出PrivacySIM评估框架,基于来自五项已发表研究的1000名用户的实证数据,评估9个前沿大模型在隐私行为模拟中的表现。这些研究涵盖医疗咨询、对话代理和聊天机器人场景。我们假设人口统计、过往经历和声明的隐私态度是隐私决策的潜在预测因子,将模型分别以这些特征子集为条件,测量其在数据共享情境下的响应与真实用户行为的一致性。结果表明:(1) 引入隐私人物画像可显著提升模拟质量,但即便最优模型准确率也仅为40.4%,仍远低于真实水平;(2) 声明的隐私态度常与实际行为不符,难以作为可靠预测指标;(3) 具有高人工智能/聊天机器人使用经验但声明隐私态度较低的用户最难以模拟。PrivacySIM为评估和改进大模型隐私模拟能力提供了首个基准,相关数据与代码已开源。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whether a core set of user persona attributes can drive LLMs to simulate individual-level privacy behavior. We introduce PrivacySIM, an evaluation suite that benchmarks LLM simulation of user privacy behavior against the ground-truth responses of 1,000 users. These users are drawn from five published user studies on privacy spanning LLM healthcare consultations, conversational agents, and chatbots. Drawing on these user studies, we hypothesize three persona facets as plausible predictors of privacy decision-making: demographics, previous experiences, and stated privacy attitudes. We condition nine frontier LLMs on subsets of these three facets and measure how often each model's response to a data-sharing scenario matches the user's actual response. Our findings show that (1) privacy persona conditioning consistently improves simulation quality over no-persona conditioning, but even the strongest model (40.4\% accuracy) remains far from faithfully simulating individual privacy decisions. (2) A user's stated privacy attitudes alone may not be the best predictor because they often diverge from the user's actual privacy behavior. (3) Users with high AI/chatbot experience but low stated privacy attitudes are the most challenging to simulate. PrivacySIM is a first step toward understanding and improving the capabilities of LLMs to simulate user privacy decisions. We release PrivacySIM to enable further evaluation of LLM privacy simulation.

隐私模拟大模型评估用户行为测评基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。