arXiv:2410.22182cs.IR2024-10中稿 · WI-IAT '24被引 10

用大模型生成个性化问答数据,降低标注成本。

Synthetic Data Generation with Large Language Models for Personalized Community Question Answering

  • 用提示工程和大模型生成用户定制化问答数据
  • 合成数据可替代人工数据训练检索系统,效果接近真实数据
  • 适合研究低资源场景下的个性化信息检索

个性化信息检索(PIR)长期受到缺乏大规模评估数据集的制约,主要因收集高质量用户信息成本高且涉及隐私。为此,本文探索利用大语言模型(LLMs)生成合成数据,以解决低资源任务的数据瓶颈。我们基于现有数据集SE-PQA,通过不同提示策略和大模型生成合成答案,构建新数据集Sy-SE-PQA。实验表明,尽管生成内容可能存在错误,但其在微调信息检索模型时表现良好,可有效替代真实人工标注数据,具备实际应用潜力。

原文摘要 · Abstract (English)

Personalization in Information Retrieval (IR) is a topic studied by the research community since a long time. However, there is still a lack of datasets to conduct large-scale evaluations of personalized IR; this is mainly due to the fact that collecting and curating high-quality user-related information requires significant costs and time investment. Furthermore, the creation of datasets for Personalized IR (PIR) tasks is affected by both privacy concerns and the need for accurate user-related data, which are often not publicly available. Recently, researchers have started to explore the use of Large Language Models (LLMs) to generate synthetic datasets, which is a possible solution to generate data for low-resource tasks. In this paper, we investigate the potential of Large Language Models (LLMs) for generating synthetic documents to train an IR system for a Personalized Community Question Answering task. To study the effectiveness of IR models fine-tuned on LLM-generated data, we introduce a new dataset, named Sy-SE-PQA. We build Sy-SE-PQA based on an existing dataset, SE-PQA, which consists of questions and answers posted on the popular StackExchange communities. Starting from questions in SE-PQA, we generate synthetic answers using different prompt techniques and LLMs. Our findings suggest that LLMs have high potential in generating data tailored to users' needs. The synthetic data can replace human-written training data, even if the generated data may contain incorrect information.

大模型生成个性化检索合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。