arXiv:2608.03700cs.CRcs.CL2026-08被引 1

研究个性化技能泄露风险,提出评估框架AntiSkillBench

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

论文配图:When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
图 1 · 摘自论文原文
  • 构建7500条对话数据集,评估技能化个人特征的泄露风险
  • 三种策略下隐私泄露与行为模仿普遍发生,攻击成功率超60%
  • 现有防御手段效果有限,难以跨策略通用

Persona skills 将个人交互历史提炼为可移植、可执行的智能体组件。虽提升个性化灵活性,但集中分散的个人信号,通过重复使用放大其影响,挑战针对单个记录或检索记忆的传统防御。为系统评估该流程的安全性,我们提出 AntiSkillBench——一个端到端基准测试框架,包含:(i) 由50个行为丰富的人物画像生成的7500条人物相关对话数据;(ii) 评估三种技能蒸馏策略下的技能级隐私泄露、代理级属性披露与行为模仿的测评套件;(iii) 覆盖在线与离线干预的四种防御配置,包括主动风险抑制与被动溯源保护。在三类前沿代理上实验表明,无论模型架构或蒸馏方式,人格技能风险持续存在,从显式属性延伸至沟通风格与人格特质。现有防御效果有限且依赖蒸馏策略,无法泛化。结果凸显 AntiSkillBench 作为隐私保护与真实性感知人格技能开发的挑战性基准。

原文摘要 · Abstract (English)

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.

隐私安全人格建模对抗攻击防御评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。