arXiv:2605.09823cs.MAcs.AI2026-05被引 4

评测多智能体大模型在日程协调中的隐私与协作权衡

CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs

论文配图:CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
图 1 · 摘自论文原文
  • 设计受控基准测试,模拟多方私有日程协同安排
  • 发现仅完成任务不等于高效:存在可避免的调度成本
  • 揭示沉默虽保隐私却可能影响团队公平负担分配

个人智能助手正作为代理处理日历、收件箱和用户偏好。日程安排使信任问题具体化:助手需与其他助手协调,同时决定向对方透露多少个人信息。我们提出CalBench,一个在私有信息约束下评估多智能体日程调度的可控基准。每个任务中,N个智能体管理各自的私有日历,调度一连串M个待办会议,以最小化干扰成本。由于无法查看他人日历,成功依赖语言协调而非集中规划。CalBench生成可解场景,配有CP-SAT最优解和非大模型参考协议,支持对任务成功率、额外成本、通信效率、负担公平性和隐私泄露的评估。在七个模型族中,我们发现仅完成任务会遗漏关键失败:智能体留下可避免的成本,通信量不能预测更低遗憾,且隐私保护性沉默可能剥夺队友所需的成本信息以实现公平分担。CalBench为研究自主助手能否在大规模部署前代表用户有效协作提供可复现的测试平台。

原文摘要 · Abstract (English)

Personal AI assistants are beginning to act as delegates with access to calendars, inboxes, and user preferences. Calendar scheduling makes the trust problem concrete: an assistant must coordinate with other assistants while deciding what to reveal about the person it represents. We introduce CalBench, a controlled benchmark for multi-agent calendar scheduling under private information. In each task, $N$ agents manage separate private calendars and schedule a stream of $M$ incoming meetings while minimizing disruption costs. Because no agent can inspect another agent's calendar, success requires language-mediated coordination rather than centralized planning. CalBench generates solvable scenarios with CP-SAT oracle solutions and decentralized non-LLM reference protocols, enabling evaluation of task success, excess cost, communication efficiency, burden fairness, and privacy leakage under matched information constraints. Across seven model families, we find that completion alone misses important failures: agents leave avoidable cost on the table, communication volume does not predict lower regret, and privacy-preserving silence can deprive teammates of cost information needed for fair burden allocation. CalBench provides a reproducible testbed for studying whether autonomous assistants can coordinate on behalf of users before deployment at scale.

多智能体日程调度隐私保护基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。