arXiv:2607.02814cs.MAcs.AI2026-07

评估用户自有的代理在隐私与合规约束下的谈判表现,发现单纯追求成功谈判不可靠。

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure

  • 构建多轮对话基准,模拟真实场景中的隐私、证据和机构压力
  • 最强基线达成率高但用户隐私泄露严重,损害实际利益
  • 新方法兼顾用户权益与可审计性,更适合真实个人代理应用

个人代理将越来越多地代表用户进行协商:分摊费用、申诉平台决定、升级支持纠纷、申请退款、更改订阅、谈判截止日期或赔偿。现有基准侧重于达成协议、增加收益或策略能力,但用户拥有的代理可能在达成协议的同时损害用户利益,如隐私泄露、同意违规、无证据主张、过度让步、升级失败或缺乏可审计性。我们提出SovereignNegotiation-Bench,一个基于追踪级别的多轮对话基准,用于评估在私有效用、披露限制、证据要求和机构不对称条件下的委托式个人代理谈判。该基准将代理可见状态与评估者独有标签分离,联合评估协议成功率、用户效用、隐私保护、同意遵守、证据依据、让步纪律、升级能力与可审计性。通过240个场景、4类模型、14个基线、13,440条冻结提示的实时轨迹、61,135行解析动作数据,以及对300项的三评审员盲审验证。最强的以协议为导向的基线虽达成率最高,但用户效用低且存在严重的隐私与同意风险;FullSovereign不追求最高协议率,却在维护用户效用、最小化信息泄露、证据支撑主张和减少未经授权承诺方面表现最佳,获得最高主权协商得分。结果表明,仅追求协议成功不足以衡量用户拥有代理的性能。

原文摘要 · Abstract (English)

Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, escalating support disputes, requesting refunds, changing subscriptions, and negotiating deadlines or reimbursements. Existing negotiation benchmarks emphasize agreement, surplus, or strategic competence, but a user-owned agent can reach an agreement while harming the user through privacy leakage, consent violation, unsupported advocacy, over-concession, failed escalation, or poor auditability. We introduce SovereignNegotiation-Bench, a trace-level multi-turn benchmark for delegated personal-agent negotiation under private utilities, disclosure constraints, evidence requirements, and institutional asymmetry. The benchmark separates agent-visible observable state from evaluator-only labels and evaluates agreement success jointly with user utility, privacy, consent, evidence grounding, concession discipline, escalation, and auditability. We report an artifact-backed validation over 240 scenarios, 4 model families, 14 baselines, 13,440 frozen-prompt live trajectories, 61,135 parsed action rows, and a blinded 3-annotator audit over 300 items. The strongest agreement-maximizing baseline achieves the highest agreement rate but low user utility and high privacy/consent risk; FullSovereign does not maximize agreement, but obtains the best sovereign negotiation score by preserving utility, minimizing leakage, grounding claims, and reducing unauthorized commitments. The results show that agreement success is insufficient for user-owned negotiation agents.

人机谈判隐私保护代理系统评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。