arXiv:2410.05252cs.CLcs.AI2024-10中稿 · EMNLP被引 2

从新闻中自动提取经济因果短叙事,助力社科研究。

Causal Micro-Narratives

  • 基于领域本体构建因果短叙事分类框架。
  • 最佳模型在检测与分类任务上分别达0.87与0.71的F1值。
  • 适用于政策分析、历史文本研究等社科场景。

我们提出一种新方法,用于从文本中分类因果微叙事。这些叙事是针对特定主题的因果或结果的句子级解释。该方法仅需主题相关的因果本体,我们在通货膨胀叙事上进行了应用。利用涵盖历史与当代美国新闻文章的人工标注数据集进行训练,评估多个大语言模型在多标签分类任务上的表现。表现最佳的模型——微调后的Llama 3.1 8B——在叙事检测任务上达到0.87的F1值,在叙事分类任务上为0.71。全面的错误分析揭示了语言歧义带来的挑战,并指出模型错误常与人工标注者分歧一致。该研究建立了一个从现实数据中提取因果微叙事的框架,具有广泛的社会科学应用前景。

原文摘要 · Abstract (English)

We present a novel approach to classify causal micro-narratives from text. These narratives are sentence-level explanations of the cause(s) and/or effect(s) of a target subject. The approach requires only a subject-specific ontology of causes and effects, and we demonstrate it with an application to inflation narratives. Using a human-annotated dataset spanning historical and contemporary US news articles for training, we evaluate several large language models (LLMs) on this multi-label classification task. The best-performing model--a fine-tuned Llama 3.1 8B--achieves F1 scores of 0.87 on narrative detection and 0.71 on narrative classification. Comprehensive error analysis reveals challenges arising from linguistic ambiguity and highlights how model errors often mirror human annotator disagreements. This research establishes a framework for extracting causal micro-narratives from real-world data, with wide-ranging applications to social science research.

因果推理文本分类大模型应用社会科学研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。