arXiv:2502.19178cs.IR2025-02KDD被引 9

构建用户嵌入评估基准,检验其在个性化问答中的提示效果。

UQABench: Evaluating User Embedding for Prompting LLMs in Personalized Question Answering

  • 设计三类任务评估用户嵌入:序列理解、行为预测、兴趣感知。
  • 实测多种方法在推荐任务中提升精度,验证嵌入有效性。
  • 提供标准化流程,适合研究个性化大模型的学者与工程师。

大型语言模型(LLMs)在自然语言处理中表现卓越。在推荐等实际场景中,用户日益追求个性化体验,将用户交互历史融入LLM上下文以增强个性化变得至关重要。然而,用户交互数据长度长且含噪声,直接作为文本提示存在挑战。一种可行方案是将交互信息压缩为紧凑的嵌入表示,作为软提示辅助LLM生成个性化回复。尽管该方法提升效率,关键问题是:用户嵌入能否充分捕捉有用信息并有效提示LLM?为此,我们提出 ame,一个用于评估用户嵌入在个性化问答中提示能力的基准。我们建立公平、标准化的评估流程,涵盖预训练、微调和评估阶段。为全面评估,设计三个维度的任务:序列理解、行为预测、兴趣感知。这些任务覆盖传统推荐任务的需求(如提升预测准确率)及基于LLM方法的愿景(如精准理解用户兴趣、改善体验)。我们在多种先进用户嵌入建模方法上进行广泛实验,并揭示了利用用户嵌入提示LLM的缩放规律。该基准已公开可用。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve remarkable success in natural language processing (NLP). In practical scenarios like recommendations, as users increasingly seek personalized experiences, it becomes crucial to incorporate user interaction history into the context of LLMs to enhance personalization. However, from a practical utility perspective, user interactions' extensive length and noise present challenges when used directly as text prompts. A promising solution is to compress and distill interactions into compact embeddings, serving as soft prompts to assist LLMs in generating personalized responses. Although this approach brings efficiency, a critical concern emerges: Can user embeddings adequately capture valuable information and prompt LLMs? To address this concern, we propose \name, a benchmark designed to evaluate the effectiveness of user embeddings in prompting LLMs for personalization. We establish a fair and standardized evaluation process, encompassing pre-training, fine-tuning, and evaluation stages. To thoroughly evaluate user embeddings, we design three dimensions of tasks: sequence understanding, action prediction, and interest perception. These evaluation tasks cover the industry's demands in traditional recommendation tasks, such as improving prediction accuracy, and its aspirations for LLM-based methods, such as accurately understanding user interests and enhancing the user experience. We conduct extensive experiments on various state-of-the-art methods for modeling user embeddings. Additionally, we reveal the scaling laws of leveraging user embeddings to prompt LLMs. The benchmark is available online.

个性化用户嵌入大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。