arXiv:2605.10606cs.CLcs.AI2026-05

研究法语中嵌入向量如何保留作者风格,发现大模型重写后仍能识别原作者特征。

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

论文配图:Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
图 1 · 摘自论文原文
  • 用文学语料对比人工写作与大模型重写,分析嵌入向量的分散度变化。
  • 嵌入向量能可靠捕捉作者风格,重写后仍保持显著差异性。
  • 适用于检测大模型伪造文本的作者身份,对版权与真实性有实用价值。

大型语言模型(LLMs)可逼真模仿人类写作风格,但尚不清楚语言模型的嵌入向量中编码了多少风格信息,以及在大模型重写后这些信息是否仍被保留。本文以法语为研究对象,采用受控的文学数据集,通过嵌入向量的分散度变化来量化风格差异的影响。实验表明,嵌入向量能可靠捕捉作者的风格特征,且这些信号在大模型重写后依然存在,并呈现出模型特有的模式。该分析结果为大模型时代下的作者身份识别提供了有前景的方向。

原文摘要 · Abstract (English)

Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embeddings from any language model and retained after LLM rewriting. We investigate these questions in French, using a controlled literary dataset to quantify the effect of stylistic variation via changes in embedding dispersion. We observe that embeddings reliably capture authorial stylistic features and that these signals persist after rewriting, while also exhibiting LLM-specific patterns. These analytical results offer promising directions for authorship imitation detection in the era of language models.

风格识别嵌入分析大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。