arXiv:2608.27818cs.AI2026-08

构建新基准评估用户与智能体在动态偏好下的协作能力。

AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics

论文配图:AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
图 1 · 摘自论文原文
  • 设计多场景动态偏好任务,模拟真实交互中的偏好演化。
  • 前沿大模型虽能处理初始模糊需求,但难以应对中途变化的偏好。
  • 仅靠提示工程无法激发模型对不确定性的识别能力。

用户在与智能体协作时,偏好很少是静态且事先明确的:偏好在交互过程中逐步形成、揭示、调整和放松。现有评估基准几乎只关注解决初始不明确的需求,难以捕捉真实交互中的复杂动态。本文提出AcCoRD,一个面向在线购物和旅行规划两个领域的用户-智能体协作评估基准,要求智能体应对多样化的用户偏好动态。我们评估了五种前沿大语言模型在两种提示策略下的表现:普通ReAct和一种基于不确定性的引导变体,该变体促使模型识别并解决用户偏好的模糊性。结果表明,尽管前沿模型能处理初始不明确的需求,但在应对交互中出现或演化的偏好时仍表现不佳,需要更复杂的不确定性建模能力。此外,仅靠提示策略无法有效激发模型对不确定性的识别。我们已开源AcCoRD,以支持开发能应对真实世界复杂偏好的智能体。

原文摘要 · Abstract (English)

User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively on resolving underspecified preferences, thereby failing to capture the richer dynamics of real-world interaction. We introduce AcCoRD, a user-agent collaboration benchmark requiring agents to handle diverse user preference dynamics in two domains: online shopping and travel planning. We evaluate five frontier LLMs under two prompting strategies: vanilla ReAct, and an uncertainty-guided variant that prompts models to identify and resolve ambiguity about user preferences. Our results reveal that frontier models can handle underspecification but struggle to satisfy preferences that emerge or evolve mid-interaction and require more sophisticated uncertainty modeling. Further, prompting alone fails to elicit the required uncertainty recognition. We release AcCoRD as a resource for developing agents that can navigate the full complexity of real-world user preferences.

人机协作动态偏好大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。