arXiv:2606.26062cs.CL2026-06被引 1

关键词词典会误判语气,让肯定句被当成自信表达

When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance

论文配图:When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance
图 1 · 摘自论文原文
  • 用大模型语义分类替代关键词计数,发现原结果是测量工具造假
  • 关键词法误把否定句中的强调词当自信,相关性从0.72降至0.21
  • 适合研究话语分析、批判性评估算法测量效度的学者阅读

计算社会科学中一个统计显著、效应量大的发现,是否可能完全源于测量工具的缺陷?我们分析了四位公共知识分子(2016–2026)共85场访谈,发现基于关键词计数的负面情绪与强调性确定性词汇存在强共现(相关系数r = 0.72–0.93,p < 0.01)。但改用大语言模型进行零样本语义分类后,该相关性大幅下降:达里奥的r从0.851降至0.206,两人呈现负相关,一人无相关。相反,大模型揭示出强烈的负面情绪与模糊表达(hedging)正相关——罗戈夫r=0.875(p=0.001),齐汉恩r=0.722(p=0.008),符合悲观言论常伴犹豫的常识。句级错误分析指出关键词词典存在三大结构性缺陷:句法盲视、多义盲视、类别缺失,导致如“从未绝对完全有信心”这类否定句被误判为高确定性。我们认为,关键词词典捕捉的是负面话语天然吸引强调性词汇的共现倾向,与修辞立场无关,甚至可能系统颠倒其含义。将关键词计数当作信念确定性的测量,是一种范畴错误。

原文摘要 · Abstract (English)

Can a statistically significant, large-effect-size finding in computational social science be entirely an artifact of the measurement instrument? We present a case where the answer appears to be yes. Analyzing 85 interviews across four public intellectuals (2016--2026), we find a robust negative-affect/emphatic-certainty lexical co-occurrence pattern under keyword-based scoring ($r = 0.72$--$0.93$, $p < 0.01$ for all four speakers). Replacing keyword counting with LLM-based zero-shot semantic classification on the complete diarized corpus (32,625 sentences) dramatically reduces this correlation: Dalio's $r = 0.851$ drops to $r = 0.206$, with two speakers showing negative $r(\text{neg}, \text{emphatic})$ and one showing null. In contrast, the LLM reveals a strong negative-hedging coupling across speakers -- Rogoff's $r(\text{neg}, \text{hedged}) = 0.875$ ($p = 0.001$) and Zeihan's $r(\text{neg}, \text{hedged}) = 0.722$ ($p = 0.008$) -- consistent with the conventional expectation that pessimistic discourse attracts hedging, not certainty. Sentence-level error analysis traces this discrepancy to three structural failure modes in keyword lexicons -- syntactic blindness, polysemy blindness, and categorical absence -- illustrated through cases where keyword counting inverts semantic meaning (e.g., ''never absolutely totally confident'' scored as high-certainty). We argue that keyword lexicons measure a universal lexical co-occurrence tendency -- negative discourse naturally attracts emphatic vocabulary -- that is orthogonal to, and can systematically invert, rhetorical stance. Treating keyword counts as measurements of epistemic certainty is a category error: a finding that appears to be about a speaker's psychology may be entirely about the counting of words.

话语分析大模型测量误差语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。