arXiv:2508.01674cs.CLcs.AI2025-08中稿 · COLM被引 22

测试大模型从对话中理解用户动态偏好能力,发现现有模型表现不佳。

CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions

  • 基于756段真实用户与模型的交互记录构建评估基准
  • 顶尖模型在推理上下文偏好时准确率低于50%,召回率不足65%
  • 适合关注个性化对话、模型行为可解释性的研究者

大型语言模型的个性化通常假设用户偏好是静态且全局一致的,但现实中偏好会随情境动态变化。当用户在不同场景中与模型互动时,会通过多轮反馈自然展现其上下文偏好,模型需从中推断并应用于未来场景以实现对齐。为此,我们提出CUPID,一个包含756条人工标注的用户-模型交互会话历史的基准。每个会话中,用户在特定情境下提出请求,并通过多轮反馈表达偏好。给定新请求和过往会话,该基准评估模型能否推断出相关偏好并生成符合偏好的回应。我们在10个开源及专有大模型上进行了测试,发现当前最先进的模型难以从多轮交互中准确推断偏好,且无法识别哪些历史上下文与新请求相关——准确率低于50%,召回率低于65%。本工作揭示了提升模型上下文感知个性化能力的迫切需求,并将CUPID作为推动改进的重要资源。

原文摘要 · Abstract (English)

Personalization of Large Language Models (LLMs) often assumes users hold static preferences that reflect globally in all tasks. In reality, humans hold dynamic preferences that change depending on the context. As users interact with an LLM in various contexts, they naturally reveal their contextual preferences, which a model must infer and apply in future contexts to ensure alignment. To assess this, we introduce CUPID, a benchmark of 756 human-curated interaction session histories between users and LLM-based chat assistants. In each interaction session, the user provides a request in a specific context and expresses their preference through multi-turn feedback. Given a new user request and prior interaction sessions, our benchmark assesses whether LLMs can infer the preference relevant to this request and generate a response that satisfies this preference. With CUPID, we evaluated 10 open and proprietary LLMs, revealing that state-of-the-art LLMs struggle to infer preferences from multi-turn interactions and fail to discern what previous context is relevant to a new request -- under 50% precision and 65% recall. Our work highlights the need to advance LLM capabilities for more contextually personalized interactions and proposes CUPID as a resource to drive these improvements.

个性化上下文理解评测基准对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。