对比了新闻灾害数据采集的自上而下与自下而上两种方法。
The Course of News Events: A Comparison of Bottom-Up and Top-Down Approaches for Collecting Text-Based Data about Disasters

- 基于时空特征用NLP聚类实现自下而上的事件发现
- 自上而下方法覆盖更全,但依赖已有灾害清单
- 方法选择影响媒体不平等研究与灾情监测结果
新闻文章是了解灾害影响与适应措施的重要信息来源。社会-环境研究中的一个关键方法论挑战是如何选取具有代表性的数据样本。常用方法有两种:一种是借助现有灾害清单,自上而下查询新闻数据库;另一种是利用自然语言处理技术,基于时间和空间特征对新闻文本进行聚类,实现自下而上的事件识别。本文以全球范围内的德国新闻报道滑坡事件数据集为例,比较了这两种方法在事件覆盖范围上的差异。研究发现,方法选择会影响最终新闻样本的构成,进而影响媒体覆盖不平等、灾害监测及灾害清单扩充等研究的应用效果。
原文摘要 · Abstract (English)
News articles are an important source of information on disaster impacts and adaptation. A key methodological challenge in socio-environmental studies is how to select a representative data sample. Two approaches are common: querying news databases top-down with the aid of an existing disaster inventory or using NLP methods to cluster news texts bottom-up based on temporal and spatial features. Using a dataset of German news about landslides worldwide, we compare these approaches and discuss variations in event coverage. Such research design decision can influence the resulting news sample, affecting its use in studies of inequality in media coverage, disaster monitoring and inventory enrichment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。