用真人隐私判断训练大模型,让代理更懂什么该说、不该说。
PrivacyAlign: Contextual Privacy Alignment for LLM Agents

- 基于599人对1350个场景的隐私标注,构建真实隐私判断数据集
- 通过标注条件奖励建模,使小模型在隐私任务上显著优于基准
- 适合关注智能代理可信性与隐私安全的研究者和开发者
代表用户决策的AI代理需符合用户真实意愿,而隐私是其中关键问题:每一次消息、发布或工具调用都涉及对分享内容、对象及条件的判断。这类判断依赖社会规范,人类不仅识别隐私违规,也定义其边界。现有方法依赖不可靠代理,本文将人类判断置于核心,提出PrivacyAlign数据集,包含1,350个样本、3,516条来自599名不同参与者的详细标注,覆盖当前大模型实际泄露隐私的多样场景。利用这些标注,我们首先证明:在参考响应上加入人类注释与解释,可提升模型判断可靠性;随后提出注释条件奖励建模,在强化学习中使用这些标注评分新响应,使小型开源模型在隐私对齐上显著改进,于PrivacyAlign与现有代理隐私基准均取得明显提升。
原文摘要 · Abstract (English)
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment problem for agents: every message, post, or tool call an agent makes is a contextual judgment about what is appropriate to share, with whom, and under which conditions. Because such judgments depend on social expectations and norms, human judgment does not merely label privacy violations but also helps define them. While existing work relies on unreliable proxies for both training and evaluation, we place human judgment at the center of agentic privacy alignment. We introduce PrivacyAlign, a dataset of 1,350 samples with 3,516 detailed annotations from 599 unique annotators across diverse scenarios where current LLMs actually leak, and use it to ground both alignment training and automated evaluation in human privacy norms. Building on these annotations, we first show that conditioning LLM judges on human annotations and explanations for reference responses to the same prompt makes their judgments more reliable. We then introduce annotation-conditioned reward modeling, which uses these annotations to score new responses during RL, and show that small open-weight agents trained with this reward better align with human privacy norms, with strong gains on PrivacyAlign and existing privacy benchmarks for agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。