大模型助写正导致语言风格趋同,研究揭示其机制与影响
Linguistic Monoculture in LLM-Assisted Language Use
- 构建作者与大模型共演化数学框架,分析三类互动模式
- 共享模型使语言趋同,个性化模型可保留多样风格差异
- 个体追求清晰会过度迎合,造成语言多样性损失
写作与交流日益依赖大语言模型(LLMs)进行起草、修改与润色。尽管这类辅助能提升语义清晰度并符合机构期待,但对统一模型的广泛依赖可能降低整体语言形式的多样性,即‘语言单一化’。我们建立数学框架,将作者与模型视为语言特征分布,并通过重复交互实现共演化。分析三类机制:固定分布的共享模型、基于作者输出递归更新的共享模型,以及结合作者特异性与群体反馈的个性化模型。结果表明,共享模型推动作者趋同于共同规范,递归反馈改变共享规范但不减少个体间差异,个性化模型则可维持多个具有非零多样性的作者-模型均衡。进一步将一致性建模为权衡清晰度、可读性与独特风格的策略选择,在此效用模型下,个体理性行为可能导致过度从众,因未内化自身独特性对他人带来的价值,形成负外部性。单个实例中‘单一化代价’有限,但当独特性主导真实性时,该代价可无限增长。合成模拟验证了固定共享、递归反馈与个性化在长期多样性上的不同影响。
原文摘要 · Abstract (English)
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。