arXiv:2604.19139cs.CLcs.AI2026-04

分析大模型口语化表达泛滥现象,揭示其与对齐训练的深层关联。

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

  • 构建量化指标VTI,系统评估8个前沿模型的口头禅频率。
  • 发现谷歌Gemini VTI高达0.590,而DeepSeek最低为0.295。
  • 人类评估显示奉承语越多,对话越不自然,相关性达-0.87。

随着大语言模型通过强化学习人类反馈(RLHF)和宪法AI等对齐技术不断演进,一种日益显著的现象浮现:口头禅——重复且程式化的语言模式广泛存在于模型输出中。这些包括阿谀奉承的开场白(如“这是一个好问题!”)、伪共情表述(如“我完全理解你的担忧”)及高频词汇(如“深入探讨”、“图景”、“细微之处”)。本文对八个顶尖模型(GPT-5.5、Claude Opus 4.8、Gemini 3.1 Pro、Grok 4.3、Doubao-Seed-2.1-pro、Kimi K2.6、DeepSeek V4 Pro、GLM-5.2)进行系统分析,基于自定义评估框架,在中英文下覆盖10类任务、10,000个提示,生成16万条响应。引入口头禅指数VTI,量化口头禅普遍性,并分析其与奉承程度、词汇多样性及人类感知自然度的相关性。结果表明模型间差异显著:Gemini 3.1 Pro的VTI最高(0.590),DeepSeek V4 Pro最低(0.295)。进一步发现,口头禅在多轮对话中累积,在主观任务中增强,且存在跨语言差异。120人参与的人类评估证实,奉承程度与自然度呈强负相关(r = -0.87,p < 0.001)。研究凸显当前对齐范式带来的“对齐代价”,呼吁建立更真实的交互框架。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs. These range from sycophantic openers ("That's a great question!", "Awesome!") to pseudo-empathetic affirmations ("I completely understand your concern", "I'm right here to catch you") and overused vocabulary ("delve", "tapestry", "nuanced"). In this paper, we present a systematic analysis of the verbal tic phenomenon across eight state-of-the-art LLMs: GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro, Grok 4.3, Doubao-Seed-2.1-pro, Kimi K2.6, DeepSeek V4 Pro, and GLM-5.2. Utilizing a custom evaluation framework for standardized API-based evaluation, we assess 10,000 prompts across 10 task categories in both English and Chinese, yielding 160,000 model responses. We introduce the Verbal Tic Index (VTI), a composite metric quantifying tic prevalence, and analyze its correlation with sycophancy, lexical diversity, and human-perceived naturalness. Our findings reveal significant inter-model variation: Gemini 3.1 Pro exhibits the highest VTI (0.590), while DeepSeek V4 Pro achieves the lowest (0.295). We further demonstrate that verbal tics accumulate over multi-turn conversations, are amplified in subjective tasks, and show distinct cross-lingual patterns. Human evaluation (N = 120) confirms a strong inverse relationship between sycophancy and perceived naturalness (r = -0.87, p < 0.001). These results underscore the "alignment tax" of current training paradigms and highlight the urgent need for more authentic human-AI interaction frameworks.

大模型对齐口语化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。