构建高质量国际事件预测数据集,助力文本驱动的全球局势研判
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling
- 用大模型生成并专家验证的标注标签,提升数据可信度
- 在真实事件预测任务中表现优异,验证数据有效性
- 开源全链路工具与数据,推动文本事件预测研究
从新闻等文本信息中预测未来国际事件,在全球政策、战略决策和地缘政治领域具有巨大潜力。然而,现有相关数据集质量普遍有限,制约了研究进展。本文提出 WORLDREP(WORLD 关系与事件预测)数据集,利用大语言模型的高级推理能力,通过精心设计的提示工程生成高质量评分标签,并由政治科学领域专家严格验证。实验表明,该数据集在真实世界事件预测任务中表现出色,具备显著实用性。我们还公开发布数据集及完整自动化数据采集、标注与基准测试代码,旨在支持并推进文本驱动事件预测的研究发展。
原文摘要 · Abstract (English)
Predicting future international events from textual information, such as news articles, has tremendous potential for applications in global policy, strategic decision-making, and geopolitics. However, existing datasets available for this task are often limited in quality, hindering the progress of related research. In this paper, we introduce WORLDREP (WORLD Relationship and Event Prediction), a novel dataset designed to address these limitations by leveraging the advanced reasoning capabilities of large-language models (LLMs). Our dataset features high-quality scoring labels generated through advanced prompt modeling and rigorously validated by domain experts in political science. We showcase the quality and utility of WORLDREP for real-world event prediction tasks, demonstrating its effectiveness through extensive experiments and analysis. Furthermore, we publicly release our dataset along with the full automation source code for data collection, labeling, and benchmarking, aiming to support and advance research in text-based event prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。