arXiv:2511.04195cs.CLcs.MA2025-11被引 2

用计算方法检验大模型文本是否像人,发现即使调优仍难骗过人类。

Computational Turing Test Reveals Systematic Differences Between Human and AI Language

  • 构建综合指标+语言特征的计算图灵测试框架
  • 九个模型调优后仍可被轻易识别,尤其在情感表达上
  • 越大越像不成立,追求拟人常牺牲内容准确性

大型语言模型(LLMs)被广泛用于社会科学中模拟人类行为,前提是它们能生成逼真的类人文本。然而这一假设尚未经过严格验证。现有评估主要依赖人类判断——即能否区分AI与人类输出,但研究表明此类判断粗糙且不可靠。为此,本文提出一种计算图灵测试:整合基于BERT的可检测性与语义相似度,结合可解释的语言特征(如风格标记和话题模式),评估LLM在特定数据集中的类人程度。我们系统比较了九个开源模型在五种校准策略下的表现,包括微调、风格提示和上下文检索,以复现X(原推特)、Bluesky和Reddit上的用户互动。结果挑战了文献中的核心假设:即便校准后,LLM输出仍明显区别于人类文本,尤其在情感基调和情绪表达方面。指令微调模型的表现反而低于基础模型,扩大模型规模也无法提升类人度。关键发现是:优化类人度常损害语义保真度,反之亦然。本研究提供了可扩展的验证与校准框架,并警示当前模型在捕捉真实人际交流方面的局限性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in the social sciences to simulate human behavior, based on the assumption that they can generate realistic, human-like text. Yet this assumption remains largely untested. Existing validation efforts rely heavily on human-judgment-based evaluations -- testing whether humans can distinguish AI from human output -- despite evidence that such judgments are blunt and unreliable. As a result, the field lacks robust tools for assessing the realism of LLM-generated text or for calibrating models to real-world data. This paper makes two contributions. First, we introduce a computational Turing test: a validation framework that integrates aggregate metrics (BERT-based detectability and semantic similarity) with interpretable linguistic features (stylistic markers and topical patterns) to assess how closely LLMs approximate human language within a given dataset. Second, we systematically compare nine open-weight LLMs across five calibration strategies -- including fine-tuning, stylistic prompting, and context retrieval -- benchmarking their ability to reproduce user interactions on X (formerly Twitter), Bluesky, and Reddit. Our findings challenge core assumptions in the literature. Even after calibration, LLM outputs remain clearly distinguishable from human text, particularly in affective tone and emotional expression. Instruction-tuned models underperform their base counterparts, and scaling up model size does not enhance human-likeness. Crucially, we identify a trade-off: optimizing for human-likeness often comes at the cost of semantic fidelity, and vice versa. These results provide a much-needed scalable framework for validation and calibration in LLM simulations -- and offer a cautionary note about their current limitations in capturing human communication.

大模型评估类人语言文本生成语言模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。