智能体应帮用户构建偏好,而非仅追问已有偏好。
Beyond expert users: agents should help users construct preferences, not just elicit them

- 提出CoPref模型,模拟用户如何通过对话形成偏好
- 在CoShop基准上,最先进模型准确率仅56%
- 适合做个性化推荐与人机交互研究的学者
传统智能体假设用户是专家,拥有明确偏好,遇到模糊任务时只用澄清问题。我们认为此假设不现实:用户常缺乏领域知识,无法直接回答偏好问题。智能体需通过示例或解释帮助用户获取必要知识,从而形成偏好。为此,我们借鉴信息经济学中的搜索-体验-信任框架,提出CoPref模型,描述用户如何基于智能体的对话行为构建偏好。进一步,在代理推荐系统中构建了CoShop交互式评测基准,评估智能体能否帮助用户获取所需知识以明确任务。测试五种前沿模型,尽管有五轮交互,最高准确率仅为56%。失败原因并非找物品能力不足,而是互动未能有效拓展用户对自身偏好的认知。
原文摘要 · Abstract (English)
Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users often lack the domain knowledge to have completely specified preferences; if asked about their preference on some feature, the user may be unable to answer without the agent helping the user to learn some domain knowledge needed to form a preference for that feature, e.g., via examples or explanations. To formalize these principles, we draw on the Search-Experience-Credence framework from Information Economics to introduce CoPref, a model of how users construct preferences based on agent dialog actions. We then study these ideas concretely in agentic recommender systems, proposing CoShop, an interactive benchmark. In CoShop, an agent converses with and makes recommendations for a CoPref user. The agent's performance depends on whether it can help the user gain the knowledge needed to specify the task well. Evaluating five frontier models, we find that no agent exceeds 56% accuracy on CoShop despite five turns of interaction. Failures stem not from agents' ability to find items, but from how little the interaction expands what users know about what they want.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。