arXiv:2603.19849cs.CLcs.AI2026-03

用语义差异度区分人类与大模型对话,简单有效。

Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue

  • 基于语义类别分布构建轻量级可解释指标
  • 大模型对话的语义差异值显著高于人类
  • 适合用于检测生成文本或研究模型模仿缺陷

大模型是否像人一样说话?这一问题在教育、学术等多个领域备受关注。本文提出一种可解释的统计特征,用于区分人类撰写与大模型生成的对话。通过情感词典分析框架,将每段文本映射为一组主题强度得分,并定义语义差值为对话中两个主导主题强度之差,假设大模型输出具有更强的主题集中性。基于多种大模型配置生成的对话数据,与脚本对白、文学作品及在线讨论等多样人类语料对比,采用韦尔奇t检验分析语义差值分布。结果显示,人工智能生成文本的语义差值始终高于人类文本,表明其主题结构更僵化,而人类对话展现出更广泛且均衡的语义分布。该零样本指标计算成本低,可作为集成检测系统中的补充信号,不替代现有方法。研究也深化了对大模型行为模仿能力的实证理解,指出主题分布是当前模型尚未达到人类对话动态的一个可量化维度。

原文摘要 · Abstract (English)

Do LLMs talk like us? This question intrigues a multitude of scholar and it is relevant in many fields, from education to academia. This work presents an interpretable statistical feature for distinguishing human written and LLMs generated dialogue. We introduce a lightweight metric derived from semantic categories distribution. Using the Empath lexical analysis framework, each text is mapped to a set of thematic intensity scores. We define semantic delta as the difference between the two most dominant category intensities within a dialogue, hypothesizing that LLM outputs exhibit stronger thematic concentration than human discourse. To evaluate this hypothesis, conversational data were generated from multiple LLM configurations and compared against heterogeneous human corpora, including scripted dialogue, literary works, and online discussions. A Welch t-test was applied to the resulting distributions of semantic delta values. Results show that AI-generated texts consistently produce higher deltas than human texts, indicating a more rigid topics structure, whereas human dialogue displays a broader and more balanced semantic spread. Rather than replacing existing detection techniques, the proposed zero-shot metric provides a computationally inexpensive complementary signal that can be integrated into ensemble detection systems. These finding also contribute to the broader empirical understanding of LLM behavioural mimicry and suggest that thematic distribution constitutes a quantifiable dimension along which current models fall short of human conversational dynamics.

对话检测可解释性大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。