arXiv:2606.19468cs.CL2026-06被引 1

首次系统分析海量文本中的叙事结构,揭示其分布不均的规律。

Characterizing Narrative Content in Web-scale LLM Pretraining Data

论文配图:Characterizing Narrative Content in Web-scale LLM Pretraining Data
图 1 · 摘自论文原文
  • 基于叙事理论构建11维分析框架,量化代理、场景与事件
  • 在300万段文本中发现连续多维叙事结构,覆盖异质数据
  • 发现不同数据源和主题间叙事质量差异显著,现有清洗无法捕捉

尽管叙事是人类交流的基础模式,但网络规模大模型预训练语料中的叙事组成仍鲜有研究。本文对一个包含3万亿词元的开源预训练语料Dolma进行了首次细粒度叙事特征分析。基于叙事理论,设计涵盖代理、场景、事件三个核心元素的11个可解释维度框架。通过采样并标注400段代表性文本,微调并验证了基于RoBERTa的NarraBERT模型,用于细粒度叙事预测。将该模型应用于300万段文本,生成新数据集NarraDolma。研究发现:(i)叙事结构可在极异质数据中大规模测量;(ii)网络文本存在连续且多维的叙事结构;(iii)叙事特质在预训练数据源与主题间分布不均,当前数据清洗策略既未测量也未考虑此类差异。本研究框架、数据集与分析为理解大模型预训练数据中叙事分布及其对叙事推理任务的影响提供了基础。NarraDolma与NarraBERT已公开发布。

原文摘要 · Abstract (English)

The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After sampling and annotating a diverse set of 400 passages, we finetune and validate NarraBERT, a RoBERTa-based model for fine-grained narrative prediction. We apply NarraBERT to 3M passages, resulting in a new dataset, NarraDolma. We find (i) narrative structure is measurable at scale across extremely heterogeneous data, (ii) we uncover a continuous, multidimensional narrative structure underlying web text, and (iii) narrative qualities are unequally distributed across pretraining sources and topics in ways that current curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data and for studying how data composition affects narrative reasoning tasks. We publicly release NarraDolma and NarraBERT.

叙事分析大模型数据文本结构可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。