arXiv:2608.10042cs.LGcs.AI2026-08

测试大模型在用户隐私下做出个性化决策的能力

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

论文配图:UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
图 1 · 摘自论文原文
  • 从交互历史推断用户隐性偏好,识别需澄清的场景
  • 多工具协作与长程行为一致性仍是主要瓶颈
  • 适合研究个性化代理、智能助手的开发者参考

工具使用的大语言模型越来越多地被要求代表用户行事,但现有评估基准通常关注身份回忆、风格模仿、通用工具使用或响应级个性化。我们提出UserToolBench,一个用于工具使用大模型个性化决策的评估基准。该基准测试模型能否从交互历史中推断潜在用户偏好,识别需要澄清的时刻,并在信息不全时生成符合用户目标的工具调用路径。基准基于经过隐私处理的真实交互记录构建,包含10个用户画像、36组工具集、1,065轮对话、170种唯一工具,涵盖缺信息、单工具和多工具等多种任务类型。对强工具使用模型的实验表明,当前模型在个性化委派方面仍存在困难,多工具协调、缺失约束推理及长程行为一致性是主要挑战。结果表明,个性化评估应从‘输出是否像用户’转向‘模型是否为所代表用户做出正确决策’。

原文摘要 · Abstract (English)

Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.

个性化决策工具调用用户建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。