用大模型自动分析公众叙事,准确率接近专家水平。
Applying Large Language Models to Characterize Public Narratives
- 基于专家共建的编码手册,用大模型实现叙事标注自动化。
- 在8个叙事、14个代码上平均F1达0.80,接近人类专家水平。
- 可扩展至政治演讲分析,助力公共话语研究与社会治理。
公众叙事(PNs)是领导力培养和公民动员的重要工具,但其系统性分析因主观解读及专家标注成本高而困难。本文提出一种新型计算框架,利用大语言模型(LLMs)自动化完成公众叙事的定性标注。我们与领域专家共同开发编码手册,并将LLM性能与专家标注对比。结果表明,LLMs在8个叙事、14个代码上平均F1达到0.80,接近人类专家水平。随后,我们在22个故事的大规模数据集上分析了叙事框架要素的分布特征。最后,将方法拓展至政治演讲分析,为公民空间中的政治修辞提供了新视角。本研究展示了大模型辅助标注在可扩展叙事分析中的潜力,同时指出了当前局限与未来方向。
原文摘要 · Abstract (English)
Public Narratives (PNs) are key tools for leadership development and civic mobilization, yet their systematic analysis remains challenging due to their subjective interpretation and the high cost of expert annotation. In this work, we propose a novel computational framework that leverages large language models (LLMs) to automate the qualitative annotation of public narratives. Using a codebook we co-developed with subject-matter experts, we evaluate LLM performance against that of expert annotators. Our work reveals that LLMs can achieve near-human-expert performance, achieving an average F1 score of 0.80 across 8 narratives and 14 codes. We then extend our analysis to empirically explore how PN framework elements manifest across a larger dataset of 22 stories. Lastly, we extrapolate our analysis to a set of political speeches, establishing a novel lens in which to analyze political rhetoric in civic spaces. This study demonstrates the potential of LLM-assisted annotation for scalable narrative analysis and highlights key limitations and directions for future research in computational civic storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。