arXiv:2603.01220cs.CL2026-03被引 1

小说数据让大模型更会编故事,也影响了AI的表达方式

Generative AI & Fictionality: How Novels Power Large Language Models

  • 用小说训练语言模型,使其更擅长生成虚构叙事
  • 相比新闻、维基等文本,小说让AI输出更具情节性和情感色彩
  • 适合关注AI文化影响与文本训练数据的研究者

生成式模型(如ChatGPT)依赖训练数据,本质是基于海量已有文本学习的下一个词预测器。自第一代GPT以来,最受欢迎的数据集普遍包含大量小说。尽管工程师普遍认为小说语言丰富,能涵盖各类社会与交流现象,但这一信念尚未被系统检验。我们以开源模型BERT为例,研究发现大语言模型不仅吸收了小说特有的语言特征,还催生出新的社会性回应形式。随着生成式AI日益塑造当代文化,对文化生产的研究必须纳入一个新维度:计算训练数据的影响。

原文摘要 · Abstract (English)

Generative models, like the one in ChatGPT, are powered by their training data. The models are simply next-word predictors, based on patterns learned from vast amounts of pre-existing text. Since the first generation of GPT, it is striking that the most popular datasets have included substantial collections of novels. For the engineers and research scientists who build these models, there is a common belief that the language in fiction is rich enough to cover all manner of social and communicative phenomena, yet the belief has gone mostly unexamined. How does fiction shape the outputs of generative AI? Specifically, what are novels' effects relative to other forms of text, such as newspapers, Reddit, and Wikipedia? Since the 1970s, literature scholars such as Catherine Gallagher and James Phelan have developed robust and insightful accounts of how fiction operates as a form of discourse and language. Through our study of an influential open-source model (BERT), we find that LLMs leverage familiar attributes and affordances of fiction, while also fomenting new qualities and forms of social response. We argue that if contemporary culture is increasingly shaped by generative AI and machine learning, any analysis of today's various modes of cultural production must account for a relatively novel dimension: computational training data.

大模型小说数据语言生成文化影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。