评测用户自有的智能代理在隐私与同意约束下的主权能力。
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
- 构建可执行基准,模拟动态意图与平台中介下的用户主权挑战。
- 120个压力场景测试显示,完整主权框架显著提升隐私保护与合规性。
- 适合关注数字主权、隐私安全的AI研发与政策制定者参考。
个人代理正成为持久的用户所有中介:记忆偏好、过滤平台信息、调用工具并协商服务。现有基准评估工具使用、网页导航、桌面控制、个性化推荐及上下文演化,但很少检验代理是否维护用户主权——即在尊重隐私、同意、证据、用户负担和抗操纵激励的前提下推进用户当前利益。我们提出SovereignPA-Bench,一个可执行的基准,用于评估用户拥有型个人代理在动态意图、平台中介、隐私边界、同意约束、证据要求及负担权衡下的表现。该基准将代理可见的ObservableState与评价者独占的HiddenLabels分离,报告任务成功、对齐度、隐私、同意、证据、操纵、负担与可审计性等分项指标,并保留情景配对顺序以支持模型与策略比较。我们在4个模型家族与8个策略基线中评估了120个主权压力场景,生成3,840条固定提示轨迹,包含原始提示、输出、服务商响应、解析动作、可重算指标、硬编码分析、定性案例及对240项的三评阅员盲审。全主权架构在主权得分上优于直接提示、仅记忆、仅同意、仅证据、ReAct/工具使用、安全提示与裁判守护基线,同时降低隐私泄露、同意违规、过度让步与操纵捕获风险。人工审计显示对隐私与同意判断高度一致,对操纵判断一致性较低,揭示了平台说服力判断的主观边界。结果表明,个人代理评估必须从任务完成转向代表性强、知情同意、证据支撑的行为评估。
原文摘要 · Abstract (English)
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user's current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. We introduce SovereignPA-Bench, an executable benchmark for evaluating user-owned personal agents under evolving intent, platform mediation, privacy boundaries, consent constraints, evidence requirements, and burden tradeoffs. The benchmark separates agent-visible ObservableState from evaluator-only HiddenLabels, reports component metrics for task success, alignment, privacy, consent, evidence, manipulation, burden, and auditability, and preserves paired scenario ordering for model and policy comparisons. We evaluate 120 sovereignty stress scenarios across 4 model families and 8 policy baselines, yielding 3,840 frozen-prompt trajectories with raw prompts, outputs, provider-form responses, parsed actions, recomputable metrics, hard-set analyses, qualitative cases, and a blinded 3-annotator audit over 240 items. Full-sovereign scaffolding improves sovereignty score over direct, memory-only, consent-only, evidence-only, ReAct/tool-use, safety-prompt, and judge-guard baselines while reducing privacy leakage, consent violation, over-concession, and manipulation capture. Human audit shows high agreement on privacy and consent and lower agreement on manipulation, identifying the subjective frontier of platform-persuasion judgments. These results show that personal-agent evaluation must move beyond task completion toward representative, consent-aware, evidence-grounded action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。