测试大模型在隐私与任务性能间的平衡,发现小模型泄露超半数敏感信息。
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

- 设计对抗性测试框架,模拟第三方恶意探查用户隐私
- 7852个样本显示小模型泄露超50%敏感属性,大模型则保留99%以上
- 为本地部署模型提供可量化的隐私对齐改进方向
大型语言模型代理越来越多地访问用户私密数据并代表用户与第三方系统交互。用户定义哪些信息可共享、哪些必须保密,代理需在第三方系统恶意行为下仍严格遵循用户意图。我们提出POLAR-Bench(政策感知对抗性基准),由具备隐私策略的可信模型与任务对话,同时面对第三方模型的对抗性探查,目标是获取任务相关信息及受保护属性。在10个领域、7,852个样本上,通过确定性集合成员检测评估隐私与任务效用,并沿两个正交维度变化隐私策略维度与攻击策略,生成每模型5×5诊断图谱。结果揭示显著分界:当前前沿模型能保留超过99%受保护属性,而1–30B规模的小型开源模型——用户最常本地运行或私有推理的类型——表现明显更差,最弱模型泄露超过一半敏感信息。POLAR-Bench因此定位了各模型意图遵循失效点,为最关键的隐私对齐提供了着力点。
原文摘要 · Abstract (English)
LLM agents increasingly have access to private user data and act on the user's behalf when interacting with third-party systems. The user defines what may and must not be shared, and the agent must robustly follow that intent even when third-party systems behave adversarially. We introduce POLAR-Bench (Policy-aware adversarial Benchmark), in which a trusted model with a privacy policy and a task converses with a third-party model that adversarially probes for both task-relevant and protected attributes. Across 10 domains and 7,852 samples, we score privacy and utility by deterministic set-membership and vary privacy policy dimension and attack strategy along two orthogonal axes, producing a 5 times 5 diagnostic surface per model. Our results reveal a sharp split: current frontier models withhold over 99% of protected attributes, while smaller open-weight models in the 1--30B range, the class users most commonly run as their own trusted agent on-device or via private inference, score notably worse, with the weakest leaking over half. POLAR-Bench thus localizes where each model's intent-following breaks down, providing a foothold for privacy alignment where it matters most.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。