arXiv:2604.19768cs.CLcs.AI2026-04

量化大模型言论与真实认知的脱节,揭示其言辞夸张却缺乏依据的特征。

Saying More Than They Know: A Framework for Quantifying Epistemic-Rhetorical Miscalibration in Large Language Models

论文配图:Saying More Than They Know: A Framework for Quantifying Epistemic-Rhetorical Miscalibration in Large Language Models
图 1 · 摘自论文原文
  • 设计三元认知-修辞标记分类体系,通过形式意义偏离等指标量化偏差
  • 大模型修辞密度接近专家两倍,且表达犹豫的标记是人类的两倍
  • 方法可自动化部署,适合检测虚假信息或训练生成内容鉴别模型

大型语言模型(LLMs)在修辞强度上常与认知基础不匹配。本研究提出一种框架,通过三元认知-修辞标记(ERM)分类体系量化这种脱节。该体系包含形式-意义偏离度(FMD)、真实-表现认知比(GPR)和修辞手法分布熵(RDDE)。在涵盖约60万词元的225篇论辩文本中,包括人类专家、非专家及大模型生成的内容,结果显示:大模型修辞使用频率接近专家的两倍(Δ=0.95),而人类作者使用反问句(erotema)的频率超过大模型两倍;大模型输出中表现性犹豫标记密度是人类的两倍。与两类人类群体相比,大模型文本的FMD显著更高(p < 0.001, Δ=0.68),且修辞手法分布更均匀。结果符合格赖斯语用学、关联理论和布拉多姆的推论主义理论预期。标注流程可完全自动化,可用于轻量级筛查人工智能生成内容的认知偏差,也可作为大模型生成文本检测的理论驱动特征集。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit systematic miscalibration with rhetorical intensity not proportionate to epistemic grounding. This study tests this hypothesis and proposes a framework for quantifying this decoupling by designing a triadic epistemic-rhetorical marker (ERM) taxonomy. The taxonomy is operationalized through composite metrics of form-meaning divergence (FMD), genuine-to-performed epistemic ratio (GPR), and rhetorical device distribution entropy (RDDE). Applied to 225 argumentative texts spanning approximately 0.6 Million tokens across human expert, human non-expert, and LLM-generated sub-corpora, the framework identifies a consistent, model-agnostic LLM epistemic signature. LLM-generated texts produce tricolon at nearly twice the expert rate ($Δ= 0.95$), while human authors produce erotema at more than twice the LLM rate. Performed hesitancy markers appear at twice the human density in LLM output. FMD is significantly elevated in LLM texts relative to both human groups ($p < 0.001, Δ= 0.68$), and rhetorical devices are distributed significantly more uniformly across LLM documents. The findings are consistent with theoretical intuitions derived from Gricean pragmatics, Relevance Theory, and Brandomian inferentialism. The annotation pipeline is fully automatable, making it deployable as a lightweight screening tool for epistemic miscalibration in AI-generated content and as a theoretically motivated feature set for LLM-generated text detection pipelines.

大模型评估认知偏差文本检测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。