构建隐私披露边界数据集,助力大模型理解用户个性化隐私偏好。
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

- 收集169人1.48万条隐私披露决策,覆盖60种人际场景
- 仅用6个历史例子就能提升预测准确率最高11.41个百分点
- 大模型更擅长结合语境理解隐私偏好,小模型依赖简单规则
将大语言模型(LLMs)与人类隐私偏好对齐,需捕捉个体在具体情境下的披露边界,而非常规隐私规范。现有研究缺乏真实场景下细微偏好的数据支持。我们提出CIDER,包含169名用户的14,850条人工标注,形成1,650组跨60种人际沟通场景的上下文披露边界数据,每组涵盖同一角色、同一AI媒介条件下9种信息共享变体的用户决策。我们设计了一个任务:基于历史边界预测用户当前披露行为,测试不同上下文信息下的模型表现。在12个开源与专有模型中,使用6个历史样本进行上下文个性化,可使预测准确率最高提升11.41个百分点。更大模型如GPT-5.4(中等推理强度)和Claude Sonnet 4.6更能利用语义上下文理解用户特定的、情境依赖的隐私偏好,而小模型则多依赖披露粒度与可识别性等结构化启发式规则。个性化普遍提升准确率,但常伴随假阳性与假阴性率失衡,仅有Claude Sonnet 4.6实现两类错误率的均衡改善。研究揭示了推理时个性化在隐私偏好建模中的潜力与局限,确立了CIDER作为推进个性化隐私对齐的重要资源。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。