arXiv:2605.29188cs.CL2026-05

用自然实验检验企业讲话中创业精神测量方法的可靠性。

Slogans or Stance? A Label-Light Diagnostic for Entrepreneurial-Discourse Measurement on Chinese SOE Speeches

论文配图:Slogans or Stance? A Label-Light Diagnostic for Entrepreneurial-Discourse Measurement on Chinese SOE Speeches
图 1 · 摘自论文原文
  • 通过同一公司不同领导讲话对比,检验测量方法是否受说话人影响。
  • 大模型在识别领导个人语言风格上表现突出,效果优于传统方法。
  • 适合研究中国国企领导话语风格差异的学者使用。

词典法、主题模型和嵌入相似性评分广泛用于管理学与社会科学研究中,以衡量企业讲话中的“创业精神”等构念。本文不提出新模型,而是构建一种标签轻量的测量诊断方法。基于80篇中央管理国企领导人讲话的语料库,利用24对同公司不同发言人及5对同公司同发言人组合,测试方法所得文档级指数是否在公司固定条件下随发言人身份变化。结果显示:LDA方法无效(Cohen d=0.20,95% CI [-0.72, 1.20]);词典评分法达到d=0.81,中文句子编码器达d=0.65(文档向量距离约10^-3)。零样本90亿参数开源大模型(Qwen3.5:9b)将配对对比d提升至1.09(精确置换检验p1=0.034)。我们下调三项结论:黄金标准F1值反映模型自身提示规则一致性而非外部构念恢复能力;文档级风格残差化使大模型d降至0.43(p1=0.22),表明约一半效应源自领导个体语言习惯;置信度加权校准虽降低方差但牺牲Δ值,且自挖掘标语词典在消融实验中几乎无影响。研究发布包含2,190段落标注语料、170段先导文本、标语词典、两组大模型评分及评估工具包。

原文摘要 · Abstract (English)

Dictionary methods, topic models, and embedding-similarity scorers are widely used in CSS and management research to measure constructs such as "entrepreneurial spirit" in corporate speeches. We contribute a label-light measurement diagnostic for such instruments rather than a new extraction model. On a corpus of 80 speeches by leaders of centrally administered Chinese state-owned enterprises, we exploit a natural experiment of 24 same-company different-speaker pairs and 5 same-company same-speaker pairs to test whether a method's per-document indices vary with leader identity holding firm constant. LDA fails (Cohen d=0.20, 95% CI [-0.72, 1.20]); a dictionary scorer reaches d=0.81 and a Chinese sentence encoder d=0.65 on doc-vector distances of order 10^-3. A zero-shot 9B open-weight LLM (Qwen3.5:9b) raises paired-contrast d to 1.09 (exact permutation p1=0.034). We downgrade three claims accordingly: gold F1 measures consistency with the LLM's own prompt rule rather than external construct recovery; doc-level style residualisation cuts the LLM's d to 0.43 (p1=0.22), so roughly half of the effect is consistent with leader idiolect; and a confidence-weighted calibration trades Delta for variance with an auto-mined slogan lexicon near-inert in ablation. We release the 2,190-segment scored corpus, the 170-paragraph pilot, the slogan lexicon, two-family LLM scores, and the evaluation harness.

话语分析大模型国企研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。