用文本数据量化气候灾害的社会经济影响,提升灾损评估的透明度与可比性。
Assessing socio-economic climate impacts from text data
- 从新闻、社交媒体等文本中提取气候灾害社会经济影响数据
- 系统梳理方法论挑战并提出标准化建议
- 适合从事灾害风险评估与气候归因研究的学者使用
自然语言处理与大语言模型的发展,使大规模新闻、社交媒体和报告文本数据得以系统化利用,构建包含洪水、干旱、风暴及多灾种事件社会经济影响的数据集。随着文本作为数据用于影响评估的研究不断扩展,其方法学复杂性也日益增加。然而,当前研究仍零散分布,缺乏对“影响”定义、时空偏差处理以及建模与后处理策略选择的明确指南,限制了不同研究间的透明度与可比性。本文通过整合常见实践,描述文本作为数据方法在分析社会经济影响时面临的关键挑战,并提出相应改进建议。通过提供最佳实践指导,旨在支持构建更稳健的文本衍生社会经济影响数据集,从而更准确地服务于灾害风险管理与归因研究。
原文摘要 · Abstract (English)
Recent advances in natural language processing (NLP) and large language models (LLMs) have enabled the systematic use of large-scale textual data from news, social media, and reports to create datasets with socio-economic impacts of climate hazards such as floods, droughts, storms, and multi-hazard events. As the field of text-as-data for impact assessment expands, so does its methodological complexity. Yet research remains fragmented, with no clear guidelines for defining what constitutes an impact, handling temporal and spatial biases, and selecting appropriate modeling and post-processing strategies. This lack of coherence limits transparency and comparability across studies. Here, we address this gap by synthesising common practices, describing key challenges specific to the use of text-as-data methods for analyzing socio-economic impact data, and proposing recommendations to address them. By providing guidance on best practices, we aim to support the construction of robust text-derived socio-economic impact datasets that can more accurately inform disaster risk management and attribution studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。