语言模型能从提示中捕捉文学风格特征
What's in a prompt? Language models encode literary style in prompt embeddings
- 用小说片段分析嵌入向量中的风格信息
- 同一作者作品的嵌入更紧密聚集,跨作者则分离
- 适合对文本风格、作者溯源感兴趣的读者
大型语言模型通过高维潜在空间编码和处理文本信息。现有研究多关注词汇概念内容如何转化为向量间的几何关系,但较少探讨整个提示的累积信息如何在变压器层作用下浓缩为单个嵌入。本文使用文学作品证明,深层表示中包含的是提示的无形而非事实性特征。我们发现,不同小说的短片段(10-100词)在潜在空间中独立分离,与它们后续预测的下一个词无关。同一作者的作品集合嵌入更紧密纠缠,跨作者则明显分散,表明嵌入编码了风格特征。这种风格几何结构可应用于作者归属与文学分析,更重要的是揭示了语言模型在信息处理与压缩上的高度复杂性。
原文摘要 · Abstract (English)
Large language models use high-dimensional latent spaces to encode and process textual information. Much work has investigated how the conceptual content of words translates into geometrical relationships between their vector representations. Fewer studies analyze how the cumulative information of an entire prompt becomes condensed into individual embeddings under the action of transformer layers. We use literary pieces to show that information about intangible, rather than factual, aspects of the prompt are contained in deep representations. We observe that short excerpts (10 - 100 tokens) from different novels separate in the latent space independently from what next-token prediction they converge towards. Ensembles from books from the same authors are much more entangled than across authors, suggesting that embeddings encode stylistic features. This geometry of style may have applications for authorship attribution and literary analysis, but most importantly reveals the sophistication of information processing and compression accomplished by language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。