构建虚构事件数据集,研究模型如何记忆事实与原文序列。
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
- 用合成虚构文本构建数据集,模拟真实网络文本分布。
- 实验证明模型能同时记住虚构事件细节和原文片段。
- 适合研究模型记忆机制的科研人员使用。
当语言模型在文本数据上训练时,它们既学习了语言结构知识,也掌握了关于世界事实的知识。在推理阶段,这些事实知识可用于解决有趣问题并为用户提供有用的信息服务。众所周知,语言模型能够逐字记忆训练数据中的长序列,但对模型如何记忆训练中出现的事实却了解甚少。本文提出一个新数据集,旨在帮助研究人员深入研究事实记忆与原文序列记忆这两种过程。该数据集包含合成生成的、类似网络文本的虚构事件文档,以及关于这些事件的问题-答案对。通过训练实验,我们展示了合成虚构事件数据在研究不同形式记忆中的有效性,并记录了构建真实感虚构合成数据所面临的挑战。
原文摘要 · Abstract (English)
When language models are trained on textual data, they acquire both knowledge about the structure of language as well as knowledge of facts about the world. At inference time, their knowledge of facts can be leveraged to solve interesting problems and perform useful knowledge work for users. It is well known that language models can verbatim memorize long sequences from their training data. However, it is much less well understood how language models memorize facts seen during training. In this work, we propose a new dataset to specifically empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events, as well as question-answer pairs about the events. We conduct training experiments showing how synthetic data about fictional events can be useful for studying different forms of memorization. We also document some challenges in effectively building realistic, fictional synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。