arXiv:2505.17048cs.CLcs.AI2025-05中稿 · NeurIPS被引 3

构建全球央行沟通数据集,用多模型分析政策信号

Words That Unite The World: A Unified Framework for Deciphering Central Bank Communications Globally

  • 构建覆盖25家央行、38万句的全球货币政策语料库
  • 跨银行联合训练模型效果显著优于单一银行模型
  • 适合研究宏观政策传播、金融自然语言处理的研究者

全球央行在维护经济稳定中发挥关键作用,解读其沟通内容的政策含义至关重要,尤其因误读可能对弱势群体造成不成比例的影响。为此,我们推出了世界央行(WCB)数据集,目前最全面的货币政策语料库,包含来自25家不同地理区域央行的超过38万句文本,覆盖28年历史数据。通过统一采样每家央行1000句(共2.5万句),并经双标注员标注、分歧解决和二次专家审核,定义了立场检测、时间分类和不确定性估计三项任务,每句均标注三类标签。我们在这些任务上对七种预训练语言模型(PLMs)和九种大语言模型(LLMs)进行基准测试(零样本、少样本及带注释引导),共完成15,075次实验。结果表明,基于跨银行聚合数据训练的模型显著优于仅使用单家银行数据训练的模型,验证了‘整体大于部分之和’的原则。通过严谨的人工评估、错误分析和预测任务,证实本框架具有经济实用性。相关资源已在HuggingFace和GitHub开源,采用CC-BY-NC-SA 4.0许可。

原文摘要 · Abstract (English)

Central banks around the world play a crucial role in maintaining economic stability. Deciphering policy implications in their communications is essential, especially as misinterpretations can disproportionately impact vulnerable populations. To address this, we introduce the World Central Banks (WCB) dataset, the most comprehensive monetary policy corpus to date, comprising over 380k sentences from 25 central banks across diverse geographic regions, spanning 28 years of historical data. After uniformly sampling 1k sentences per bank (25k total) across all available years, we annotate and review each sentence using dual annotators, disagreement resolutions, and secondary expert reviews. We define three tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation, with each sentence annotated for all three. We benchmark seven Pretrained Language Models (PLMs) and nine Large Language Models (LLMs) (Zero-Shot, Few-Shot, and with annotation guide) on these tasks, running 15,075 benchmarking experiments. We find that a model trained on aggregated data across banks significantly surpasses a model trained on an individual bank's data, confirming the principle "the whole is greater than the sum of its parts." Additionally, rigorous human evaluations, error analyses, and predictive tasks validate our framework's economic utility. Our artifacts are accessible through the HuggingFace and GitHub under the CC-BY-NC-SA 4.0 license.

央行沟通政策分析多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。