用小说中的隐私规范训练大模型,让其更懂真实场景下的隐私边界。
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction

- 从小说中提取隐私规范结构,用于监督微调和强化学习。
- 在法律合规任务上得分最高,且与人类隐私判断相关性最强。
- 适合研究隐私对齐、伦理推理或可解释AI的开发者使用。
大型语言模型代理的信息处理行为普遍与用户的情境化隐私期望不符。情境完整性(CI)提供了一个原则性框架,将隐私定义为在特定情境下信息流动的适当性。现有方法要么通过监督-助手架构增加推理开销,要么在窄任务数据上微调。本文提出从虚构小说中提取规范拟像(即规范与信息流的结构化表示),并采用监督学习微调后接GRPO强化学习进行训练。复合奖励函数结合程序化信号(包括任务清晰度、结构完整性、内部一致性及上下文识别)与大模型裁判,评估模型推理是否基于源文本的隐含规范体系。为防止过拟合,引入每完成项对比评分:每个输出同时与正确规范体系和随机错误体系比较,促使模型依赖上下文而非记忆源特定规范。我们在五个涵盖不同社会情境的CI对齐基准上评估,剖析了强化学习与规范基底的贡献。在七种模型中,监督微调引入保守先验,增强对隐私相关情境的识别,但未提升判断正确性;而结合规范基底的GRPO达到法律合规基准最高分,并与众包人类隐私预期相关性最强,证明小说衍生的规范拟像可有效传授跨现实领域的上下文隐私推理能力。
原文摘要 · Abstract (English)
Information handling practices of LLM agents are broadly misaligned with the contextual privacy expectations of their users. Contextual Integrity (CI) provides a principled framework, defining privacy as the appropriate flow of information within context-relative norms. However, existing approaches either double inference cost via supervisor-assistant architectures, or fine-tune on narrow task-specific data. We propose extracting normative simulacra (structured representations of norms and information flows) from fiction novels and using them to fine-tune LLMs via supervised learning followed by GRPO reinforcement learning. Our composite reward function combines programmatic signals, including task clarity (subsuming schema validity, construct discrimination, and extraction confidence), structural completeness, internal consistency, and context identification, with an LLM judge that evaluates whether the model's privacy reasoning is grounded in the held-out normative universe of the source text. To mitigate overfitting, we introduce per-completion contrastive scoring: each completion is evaluated against both the correct normative universe and a randomly selected wrong one, teaching the model to condition on context rather than memorize source-specific norms. We evaluate on five CI-aligned benchmarks spanning distinct societal contexts and ablate the contributions of RL and normative grounding. Across seven models, SFT introduces a conservative prior toward restricting information flow, improving recognition of privacy-relevant situations but not the correctness of privacy judgments. GRPO with normative grounding achieves the highest score on a law compliance benchmark and strongest correlation with crowdsourced human privacy expectations, demonstrating that fiction-derived normative simulacra can teach contextual privacy reasoning that transfers to real-world domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。