arXiv:2410.03905cs.CL2024-10NeurIPS被引 7

构建首个用户主观导向的摘要数据集,检验大模型生成摘要是否贴合普通人需求。

PersonalSum: A User-Subjective Guided Personalized Summarization Dataset for Large Language Models

  • 基于真实用户标注,构建个性化摘要数据集PersonalSum。
  • 发现实体主题仅是影响偏好因素之一,大模型仍难满足个体差异。
  • 适合研究个性化生成、人机交互与可解释性的人群参考。

近年来自然语言处理快速发展,大量研究表明,大语言模型(LLMs)生成的通用摘要在人类评估中有时甚至优于记者等专家标注的摘要。然而,这些通用摘要是否满足普通用户的个体化需求尚缺乏研究。主要瓶颈在于缺少公众参与的真实标注数据集。现有个性化摘要研究多依赖从通用摘要数据集中衍生的伪数据,或基于假设任务生成的可控摘要,缺乏用户主动参与。为此,我们提出了首个高质量、人工标注的个性化摘要数据集PersonalSum,首次探究公众读者关注点与大模型通用摘要之间的差异。该数据集包含用户画像、个性化摘要及其对应原文句子,以及机器生成的通用摘要及来源。我们通过少样本上下文学习场景,考察实体/主题、情节结构等个人信号对摘要生成的影响。初步结果表明,实体主题仅为影响用户偏好的关键因素之一,个性化摘要仍是当前大模型面临的重要挑战。

原文摘要 · Abstract (English)

With the rapid advancement of Natural Language Processing in recent years, numerous studies have shown that generic summaries generated by Large Language Models (LLMs) can sometimes surpass those annotated by experts, such as journalists, according to human evaluations. However, there is limited research on whether these generic summaries meet the individual needs of ordinary people. The biggest obstacle is the lack of human-annotated datasets from the general public. Existing work on personalized summarization often relies on pseudo datasets created from generic summarization datasets or controllable tasks that focus on specific named entities or other aspects, such as the length and specificity of generated summaries, collected from hypothetical tasks without the annotators' initiative. To bridge this gap, we propose a high-quality, personalized, manually annotated abstractive summarization dataset called PersonalSum. This dataset is the first to investigate whether the focus of public readers differs from the generic summaries generated by LLMs. It includes user profiles, personalized summaries accompanied by source sentences from given articles, and machine-generated generic summaries along with their sources. We investigate several personal signals - entities/topics, plot, and structure of articles - that may affect the generation of personalized summaries using LLMs in a few-shot in-context learning scenario. Our preliminary results and analysis indicate that entities/topics are merely one of the key factors that impact the diverse preferences of users, and personalized summarization remains a significant challenge for existing LLMs.

个性化摘要数据集大模型用户偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。