用大模型标注社交媒体中的价值观表达,提升标注一致性与预测准确性。
Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

- 基于舒瓦茨价值观理论,设计可复现的提示策略进行价值标注。
- 通过迭代校准降低错误率,使大模型标注更接近专家判断。
- 将模糊标注结果迁移至编码器模型,适合大规模社会情绪分析。
在自然语言社交媒体文本中测量主观概念,需具备理论基础、实证验证且可迁移至编码器模型以实现规模化预测。本文基于舒瓦茨基本人类价值观理论,对用户帖子进行标注,探究不同大模型与提示策略(即标注范式)如何在文本中操作化表达价值观。除标准分类指标外,还评估结构对齐性、标注歧义性、错误模式及重复运行稳定性。研究发现,不同大模型会产生不同的价值解读;通过错误分析进行迭代提示校准,可减少误判并提升与专家标注的一致性。错误模式进一步用于构建针对性专家验证规则。将大模型生成的软标签迁移至编码器模型进行预测,保留了价值表达中的不确定性信息。对超过一百万条帖子的敏感性分析显示,不同标注范式差异会传播至预测的价值表达水平,而标准化的时间动态和重大事件响应方向则更具鲁棒性。
原文摘要 · Abstract (English)
Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically validated, and transferable to an encoder model for scalable prediction. Using posts annotated according to Schwartz's theory of basic human values, we investigate how different LLMs and prompting strategies, which we call annotation regimes, operationalize the expression of values in text. Beyond standard classification metrics, we evaluate structural alignment, annotation ambiguity, error patterns, and stability across repeated runs. We find that different LLMs produce different value interpretations, and iterative prompt calibration through error analysis reduces misattributions and improves alignment with expert annotations. Error patterns are further used to derive targeted expert-verification rules for corpus annotation. We transfer soft LLM labels to an encoder model for prediction, retaining information about ambiguity in value expression. Finally, a sensitivity analysis on more than one million posts shows that regime-specific annotation differences propagate into predicted levels of value expression, whereas standardized temporal dynamics and the direction of major event responses are more robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。