研究大模型立场分析中预处理与测量方法的不稳定性来源。
Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse
- 区分预处理流程与大模型标注的干扰,诊断不一致根源。
- 低样本发言者(≤5段)受预处理影响大,高样本者(≥16段)影响小(均值仅0.13)。
- 不同测量方法(大模型vs词典)常得出相反结论,需重点验证方法选择。
计算社会科学研究日益依赖自动化预处理流程——说话人分离、语音识别文本清洗、句子分割——将原始媒体转化为可分析文本。当同一输入产生不同输出时,不稳定性的两个来源为:预处理流程本身(如说话人分离方法、分句规则)和下游测量工具(大模型标注与关键词词典)。我们基于41位公众人物在五个领域中的256个YouTube访谈,对比了两种说话人分离流程和两种测量方法,均聚焦情感极性与认知模态的耦合关系。结果发现:(1)预处理敏感性集中在视频样本少的发言者(N ≤ 5);对四个样本最多的发言者(N ≥ 16),预处理导致的 $r( ext{neg}, ext{emph})$ 绝对变化均值仅为0.13;(2)跨方法分歧更大且系统化——即使在同一预处理流程下,大模型与词典方法也对部分高样本发言者给出相反的耦合方向;(3)总体负面情绪比例高度稳定(|Δp(负向)| < 6个百分点),掩盖了上述两种不稳定性来源。本研究提出一种诊断框架,可分离流程效应与测量效应:研究访谈数据中跨维度关系的学者应验证其结论对两类变异的稳健性,尤其关注测量方法的选择。
原文摘要 · Abstract (English)
Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce different outputs from the same input, two distinct sources of instability can arise: the preprocessing pipeline itself (diarization method, segmentation rules) and the downstream measurement instrument (LLM annotation vs.\ keyword lexicon). Using 256 YouTube interviews across 41 public figures from five domains, we compare two speaker-diarization pipelines and two measurement methods, all targeting the coupling between affective valence and epistemic modality. We find that (1) preprocessing pipeline sensitivity is concentrated in speakers with limited video samples (N $\leq 5$); for the four best-sampled speakers (N $\geq 16$), the mean absolute pipeline-induced change in $r(\text{neg}, \text{emph})$ is only $0.13$; (2) cross-method disagreement is larger and more systematic -- the LLM and keyword-lexicon methods assign opposite coupling directions to several well-sampled speakers, even within the same preprocessing pipeline; and (3) aggregate valence proportions are highly stable ($|Δp(\text{neg})| < 6$pp) regardless of pipeline or method, masking both sources of instability. The contribution is a diagnostic framework that separates pipeline effects from measurement effects: researchers studying cross-dimensional relationships in interview data should verify that their conclusions are robust to both sources of variation, with particular attention to measurement method choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。