arXiv:2506.00137cs.CLcs.IR2025-06EMNLP被引 35

构建首个个性化长文本问答基准,提升回答贴合用户需求的能力

LaMP-QA: A Benchmark for Personalized Long-form Question Answering

  • 设计覆盖45个子类的个性化问答数据集
  • 引入用户上下文使回答性能最高提升39%
  • 适合研究个性化对话与大模型应用的学者

个性化对以用户为中心的问答系统至关重要,但相关研究仍不充分,主要因缺乏训练与评估资源。为此,我们提出LaMP-QA——一个面向个性化长文本答案生成的基准测试。该数据集涵盖艺术与娱乐、生活与个人发展、社会与文化三大类别,共45个子类。通过人工与自动评估,我们对比了多种评价策略,并衡量生成答案与人类偏好的一致性。同时,基于开源与专有大语言模型,评估了多种非个性化与个性化方法。结果表明,引入个性化上下文可使性能最高提升39%。该基准已公开发布,以支持该领域未来研究。

原文摘要 · Abstract (English)

Personalization is essential for question answering systems that are user-centric. Despite its importance, personalization in answer generation has been relatively underexplored. This is mainly due to lack of resources for training and evaluating personalized question answering systems. We address this gap by introducing LaMP-QA -- a benchmark designed for evaluating personalized long-form answer generation. The benchmark covers questions from three major categories: (1) Arts & Entertainment, (2) Lifestyle & Personal Development, and (3) Society & Culture, encompassing over 45 subcategories in total. To assess the quality and potential impact of the LaMP-QA benchmark for personalized question answering, we conduct comprehensive human and automatic evaluations, to compare multiple evaluation strategies for evaluating generated personalized responses and measure their alignment with human preferences. Furthermore, we benchmark a number of non-personalized and personalized approaches based on open-source and proprietary large language models. Our results show that incorporating the personalized context provided leads to up to 39% performance improvements. The benchmark is publicly released to support future research in this area.

问答系统个性化大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。