arXiv:2509.13539cs.CL2025-09ACL被引 1

构建首个美联储货币政策立场标注数据集,助力模型理解政策语义

Op-Fed: Opinion, Stance, and Monetary Policy Annotations on FOMC Transcripts Using Active Learning

  • 设计分层标注框架,分离观点、政策与立场三要素
  • 采用主动学习提升稀有正例数量,使正类样本翻倍
  • 发现大模型对政策立场识别准确率不足人类水平

美国联邦公开市场委员会(FOMC)定期讨论并制定货币政策,影响数百万人的借贷与支出决策。本文发布Op-Fed数据集,包含从FOMC会议记录中人工标注的1044个句子及其上下文。构建过程中面临两大技术挑战:类别极度不平衡——我们估计表达非中立政策立场的句子不足8%;以及句间依赖性——65%的实例需要超出单句的上下文。为此,我们设计了五阶段分层标注方案,以分离观点、货币政策和政策立场三个维度及所需上下文层级。其次,采用主动学习选取待标注实例,使各标注维度的正例数量大致翻倍。基于Op-Fed,我们发现性能最佳的闭源大模型在观点分类上零样本准确率达0.80,但在政策立场分类上仅0.61,低于人类基线0.89。我们预计该数据集将推动未来模型训练、置信度校准及新标注工作的开展。

原文摘要 · Abstract (English)

The U.S. Federal Open Market Committee (FOMC) regularly discusses and sets monetary policy, affecting the borrowing and spending decisions of millions of people. In this work, we release Op-Fed, a dataset of 1044 human-annotated sentences and their contexts from FOMC transcripts. We faced two major technical challenges in dataset creation: imbalanced classes -- we estimate fewer than 8% of sentences express a non-neutral stance towards monetary policy -- and inter-sentence dependence -- 65% of instances require context beyond the sentence-level. To address these challenges, we developed a five-stage hierarchical schema to isolate aspects of opinion, monetary policy, and stance towards monetary policy as well as the level of context needed. Second, we selected instances to annotate using active learning, roughly doubling the number of positive instances across all schema aspects. Using Op-Fed, we found a top-performing, closed-weight LLM achieves 0.80 zero-shot accuracy in opinion classification but only 0.61 zero-shot accuracy classifying stance towards monetary policy -- below our human baseline of 0.89. We expect Op-Fed to be useful for future model training, confidence calibration, and as a seed dataset for future annotation efforts.

自然语言处理政策分析主动学习大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。