大模型重写个人叙事时会统一风格,削弱个人语气。
Voice Under Revision: Large Language Models and the Normalization of Personal Narrative
- 用三个大模型在不同提示下重写300篇叙事,分析13个语言特征变化
- 重写后功能词、代词减少,词汇多样性和标点复杂度上升,风格趋同
- 即使强调保留声音,仍无法避免风格标准化,适合关注文本真实性研究者
本研究考察大语言模型重写如何改变个人叙事的风格与叙述质感。分析了300篇个人叙事经三类前沿LLM在三种提示条件下(通用优化、仅重写、保留语调)重写后的变化:通用改进、仅重写、保留语调。通过计算文体学中的13个语言指标(包括虚词、词汇多样性、词长、标点、缩略语、第一人称代词、情感词)进行测量。结果显示,无论模型或提示条件,重写均导致一致的风格趋同。虚词、缩略语和第一人称代词减少,词汇多样性、词长和标点复杂度增加。这种变化在要求“改进”或“重写”的提示下均发生。保留语调提示虽减小变化幅度,但不改变方向。样式分析显示,重写文本在特征空间中趋于聚集,难以回溯至原文。额外叙述标记表明,叙事从嵌入式转为疏离式,因果推理由显性转为压缩抽象。结果说明当前大模型具有将文本导向更精炼、更去情境化语域的定向作用。这对数字人文与计算文本分析有重要影响,因虚词、代词、缩略语、标点等常作为风格、语调、作者身份和语料完整性的证据。因此,大模型修订不应仅视为表面编辑,而应理解为具有实质影响的文本中介行为。
原文摘要 · Abstract (English)
This study examines how large language model rewriting alters the style and narrative texture of personal narratives. It analyzes 300 personal narratives rewritten by three frontier LLMs under three prompt conditions: generic improvement, rewrite-only, and voice-preserving revision. Change is measured across 13 linguistic markers drawn from computational stylistics, including function words, vocabulary diversity, word length, punctuation, contractions, first-person pronouns, and emotion words. Across models and prompt conditions, LLM rewriting produces a consistent pattern of stylistic normalization. Function words, contractions, and first-person pronouns decrease, while vocabulary diversity, word length, and punctuation elaboration increase. These shifts occur whether the prompt asks the model to "improve" the text or simply to "rewrite" it. Voice-preserving prompts reduce the magnitude of the changes but do not eliminate their direction. Stylometric analysis shows that rewritten texts converge in feature space and become harder to match back to their source texts. Additional narrative markers indicate a shift from embedded to distanced narration, and from explicit causal reasoning to compressed abstraction. The findings suggest that contemporary LLMs exert a directional pull toward a more polished, less situated register. This has consequences for digital humanities and computational text analysis, where features such as function words, pronouns, contractions, and punctuation often serve as evidence for style, voice, authorship, and corpus integrity. LLM revision should therefore be understood not merely as surface-level editing, but as a consequential form of textual mediation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。