BERT能编码虚构叙事中的时间、空间、因果和角色维度。
Do BERT Embeddings Encode Narrative Dimensions? A Token-Level Probing Analysis of Time, Space, Causality, and Character in Fiction

- 用线性探测分析BERT词元嵌入,验证其是否蕴含叙事语义。
- 94%准确率证明编码有效,罕见维度如因果召回达0.75。
- 发现边界泄露现象,适合研究叙事表征的学者关注。
叙事理解依赖多维语义结构。本研究探究BERT嵌入是否编码虚构叙事语义的四个维度:时间、空间、因果与角色。借助大模型加速标注,构建了基于词元级别的数据集,包含上述四类及“其他”类别标签。对BERT嵌入的线性探测达到94%准确率,显著优于方差匹配的随机嵌入控制组(47%),证实BERT编码了有意义的叙事信息。采用均衡类别加权后,宏观平均召回率达0.83,对罕见类别如因果(召回=0.75)和空间(召回=0.66)表现中等。混淆矩阵分析揭示‘边界泄漏’现象,即稀有维度常被误判为‘其他’。聚类分析显示无监督聚类与预定义类别近乎随机一致(ARI = 0.081),表明叙事维度虽被编码但未形成离散可分的簇。未来工作包括引入仅词性标注基线以分离句法模式与叙事编码、扩展数据集及逐层探测。
原文摘要 · Abstract (English)
Narrative understanding requires multidimensional semantic structures. This study investigates whether BERT embeddings encode dimensions of fictional narrative semantics -- time, space, causality, and character. Using an LLM to accelerate annotation, we construct a token-level dataset labeled with these four narrative categories plus "others." A linear probe on BERT embeddings (94% accuracy) significantly outperforms a control probe on variance-matched random embeddings (47%), confirming that BERT encodes meaningful narrative information. With balanced class weighting, the probe achieves a macro-average recall of 0.83, with moderate success on rare categories such as causality (recall = 0.75) and space (recall = 0.66). However, confusion matrix analysis reveals "Boundary Leakage," where rare dimensions are systematically misclassified as "others." Clustering analysis shows that unsupervised clustering aligns near-randomly with predefined categories (ARI = 0.081), suggesting that narrative dimensions are encoded but not as discretely separable clusters. Future work includes a POS-only baseline to disentangle syntactic patterns from narrative encoding, expanded datasets, and layer-wise probing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。