分析BERT如何分组叙事内容,发现它更关注语义而非作者风格。
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT
- 逐层分析BERT激活值,用PCA和MDS找模式
- 后期层对叙事内容聚类明显,风格聚类弱
- 适合研究模型语义表征与认知机制的读者
本研究聚焦于基于Transformer的大型语言模型BERT,探究其在不同层中对叙事内容与作者风格的聚类能力。利用GPT-4生成的包含多样语义内容与风格差异的叙事数据集,通过主成分分析(PCA)与多维缩放(MDS)技术分析BERT各层激活值,发现其后期层对叙事内容呈现强聚类特征,且聚类日益紧凑分明;当文本被重述为寓言、科幻或儿童故事等不同类型时,风格聚类较强,但针对个别作者的独特写作风格则几乎无聚类现象。结果表明,BERT更倾向于捕捉语义内容而非作者个性特征,揭示了其表征能力与层级处理机制,为理解Transformer模型的语言编码方式提供了新视角,并推动人工智能与认知神经科学的跨学科研究。
原文摘要 · Abstract (English)
This study investigates the internal mechanisms of BERT, a transformer-based large language model, with a focus on its ability to cluster narrative content and authorial style across its layers. Using a dataset of narratives developed via GPT-4, featuring diverse semantic content and stylistic variations, we analyze BERT's layerwise activations to uncover patterns of localized neural processing. Through dimensionality reduction techniques such as Principal Component Analysis (PCA) and Multidimensional Scaling (MDS), we reveal that BERT exhibits strong clustering based on narrative content in its later layers, with progressively compact and distinct clusters. While strong stylistic clustering might occur when narratives are rephrased into different text types (e.g., fables, sci-fi, kids' stories), minimal clustering is observed for authorial style specific to individual writers. These findings highlight BERT's prioritization of semantic content over stylistic features, offering insights into its representational capabilities and processing hierarchy. This study contributes to understanding how transformer models like BERT encode linguistic information, paving the way for future interdisciplinary research in artificial intelligence and cognitive neuroscience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。