arXiv:2510.08986cs.CLcs.CE2025-10ACL

构建首个中文政策文本多级语义标注数据集,助力政策语言分析与AI研究

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

  • 基于五色分类法对330万段政策文本进行专家标注
  • 跨74年覆盖1.9万份中央政策文件,标注一致性达K=0.86
  • 适合政策分析、法律NLP及多语言智能系统研究者使用

我们推出CAPC-CG——中国适应性政策沟通(中央政府)语料库,是首个公开的中文政策指令语料库,采用五色分类法对清晰与模糊语言进行标注,基于安氏适应性政策沟通理论。语料涵盖1949至2023年间中国最高权威机构发布的国家法律、行政法规及部门规章,共330万段落单位。配套发布完整元数据、双轮标注框架及由专家与训练编码员构建的黄金标准标注集。标注者间一致性达Fleiss's kappa K = 0.86,表明适用于监督建模。提供多种大语言模型的基线分类结果、标注手册,并描述数据集中的模式特征。本发布旨在支持下游任务及政策沟通领域的多语言自然语言处理研究。

原文摘要 · Abstract (English)

We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives annotated with a five-color taxonomy of clear and ambiguous language categories, building on Ang's theory of adaptive policy communication. Spanning 1949-2023, this corpus includes national laws, administrative regulations, and ministerial rules issued by China's top authorities. Each document is segmented into paragraphs, producing a total of 3.3 million units. Alongside the corpus, we release comprehensive metadata, a two-round labeling framework, and a gold-standard annotation set developed by expert and trained coders. Inter-annotator agreement achieves a Fleiss's kappa of K = 0.86 on directive labels, indicating high reliability for supervised modeling. We provide baseline classification results with several large language models (LLMs), together with our annotation codebook, and describe patterns from the dataset. This release aims to support downstream tasks and multilingual NLP research in policy communication.

政策分析语料库大模型中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。