用智能流程生成地震灾情模拟推文,替代真实数据难获取问题。
Design and evaluation of an agentic workflow for crisis-related synthetic tweet datasets
- 设计迭代式智能工作流,按目标特征生成合成推文。
- 生成的推文准确匹配指定地点与损毁等级标签。
- 适合需大规模灾情数据的AI系统评估与训练场景。
Twitter(现为X)已成为危机情境下态势感知的重要社交媒体数据来源。危机计算研究广泛使用推文开发和评估人工智能系统,如从推文中提取位置、估算损毁程度以支持灾情评估。然而,近期Twitter的数据访问政策变化使得构建真实推文灾情数据集愈发困难。此外,现有数据集仅限于特定历史事件,且大规模标注成本高昂。为解决这些问题,我们提出一种用于生成灾情相关合成推文数据集的智能工作流。该工作流通过预设目标特征迭代生成合成推文,利用预定义合规性检查评估质量,并引入结构化反馈在后续迭代中优化内容。以震后损毁评估为例,我们证明该工作流可生成准确反映目标地点与损毁等级的合成推文。进一步实验表明,生成的数据集可用于评估地理定位与损毁等级预测等任务中的AI系统。结果表明,该方法提供了一种灵活、可扩展的替代方案,能系统生成适用于多样危机事件、社会背景及应用的合成社交媒体数据。
原文摘要 · Abstract (English)
Twitter (now X) has become an important source of social media data for situational awareness during crises. Crisis informatics research has widely used tweets from Twitter to develop and evaluate artificial intelligence (AI) systems for various crisis-relevant tasks, such as extracting locations and estimating damage levels from tweets to support damage assessment. However, recent changes in Twitter's data access policies have made it increasingly difficult to curate real-world tweet datasets related to crises. Moreover, existing curated tweet datasets are limited to past crisis events in specific contexts and are costly to annotate at scale. These limitations constrain the development and evaluation of AI systems used in crisis informatics. To address these limitations, we introduce an agentic workflow for generating crisis-related synthetic tweet datasets. The workflow iteratively generates synthetic tweets conditioned on prespecified target characteristics, evaluates them using predefined compliance checks, and incorporates structured feedback to refine them in subsequent iterations. As a case study, we apply the workflow to generate synthetic tweet datasets relevant to post-earthquake damage assessment. We show that the workflow can generate synthetic tweets that capture their target labels for location and damage level. We further demonstrate that the resulting synthetic tweet datasets can be used to evaluate AI systems on damage assessment tasks like geolocalization and damage level prediction. Our results indicate that the workflow offers a flexible and scalable alternative to real-world tweet data curation, enabling the systematic generation of synthetic social media data across diverse crisis events, societal contexts, and crisis informatics applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。