用主题模型当分类器,精准筛选德语新闻中的极端气候事件报道。
Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News

- 基于主题模型后验分布筛选相关新闻,不改变训练流程。
- 在7类极端气候事件中,关键词概率提升样本精确率。
- 结果因灾害类型而异,不宜将气候事件混为一谈。
在极端气候事件媒体报道研究中,自然语言处理方法已成为从大型新闻数据库中识别相关内容的必备工具。然而,训练高精度深度学习分类器所需的标注数据往往不足。主题模型具有无监督和可解释的优势,但通常仅用于探索性分析或数据特征描述。本研究探讨如何将主题模型作为二分类器,用于德国媒体中七类极端气候事件的新闻检索。方法基于主题模型估计的后验分布,无需修改训练过程即可筛选相关文档。利用标注样本评估,发现用于查询新闻数据库的关键词概率也能有效选择相关主题,提升样本精确率。与微调文本嵌入分类器及开源大模型对比,结果显示大模型精度最低。此外,性能表现受灾害类型影响,表明不应将气候事件统一视为单一类别进行自然语言处理任务。
原文摘要 · Abstract (English)
In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. Still, enough annotated data to train accurate deep learning-based classifiers from scratch is often not available. Topic Models have the advantage of being both unsupervised and interpretable, but are typically used only for exploratory analysis or data characterisation. In this study, we investigate how to employ Topic Models as binary classifiers for refining the retrieval of relevant news about seven types of extreme climate events in the German media. Our method relies on the posterior distributions estimated by Topic Models to select relevant documents, without modifying their training procedure. Using an annotated sample to guide the evaluation, we show that the probabilities assigned to keywords used to query news databases can also be informative for selecting relevant topics and improve sample precision. We compare our results to a fine-tuned text embedding classifier and an open-weight LLM, discussing observed trade-offs, e.g. the LLM's lowest precision. Moreover, we show that results are hazard-dependent, which speaks against considering climate events as a single category in NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。