arXiv:2606.06614cs.CLcs.AI2026-06被引 1

用真人对话数据发现大模型个性化存在三重短板,提出轻量改进方案。

Re-Centering Humans in LLM Personalization

论文配图:Re-Centering Humans in LLM Personalization
图 1 · 摘自论文原文
  • 基于真人对话数据构建三阶段评估体系,真实检验个性化能力。
  • 模型在提取用户信息、匹配属性与生成回应上表现均低于人类判断。
  • 虽轻量训练可提升前两阶段表现,但难建模人类对个性化质量的判断。

尽管关注度上升,当前大语言模型(LLM)个性化评估仍主要依赖合成数据,其在真实用户场景下的表现尚不清晰。本文通过收集550段真人对话及三阶段判断数据——从对话中提取用户属性(5,949次判断)、匹配相关属性至新提示(11,919次)、生成个性化回应(1,101次),揭示了系统在各阶段的局限性。模型难以准确提取真人对话中的属性,对相关属性的判断与人类存在分歧,生成的个性化回应也被人类评价为不优于通用回复(尽管模型自身评分更高)。我们引入两种轻量级训练干预,在前两个阶段使自动化评估更贴近人类数据。但在第三阶段,学习到的奖励模型与人类评分仅呈弱相关,表明人类对个性化质量的判断难以直接建模。本研究收集的数据为理解模型如何有效提取、选择并融入用户信息提供了基础。

原文摘要 · Abstract (English)

Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains unclear how well current personalization systems work for real users. In this paper, we study the gap in LLM personalization performance when using synthetic versus human data. We collect human conversations (550 conversations) and judgments across three stages of personalization: extracting user attributes from conversations (5,949 judgments), pairing relevant attributes with new prompts (11,919), and incorporating relevant attributes into a personalized response (1,101). Incorporating human data reveals system limitations at each stage. Models struggle to extract attributes from human conversations, disagree with human judgments on relevant attributes, and generate personalized responses that humans judge no better than generic responses (though that LLM judges widely rate as better). We introduce two lightweight training-based interventions that shift automated personalization evaluation closer to human data in our first two stages. However, in our third stage we find that learned reward models achieve only modest correlation with human ratings, suggesting that human-aligned personalization quality judgments are difficult to model directly. Our collected data provides a foundation for studying how models should extract, select, and incorporate user information in ways that humans find useful.

大模型个性化人类对齐评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。