arXiv:2503.04439cs.CL2025-03被引 1

首个中文新闻框架数据集,助力自动识别中文媒体叙事倾向。

A Dataset for Analysing News Framing in Chinese Media

  • 构建首个针对中文新闻的框架分析数据集,涵盖复杂语义特征。
  • 基于XLM-RoBERTa微调获0.753的F1-micro分数,优于零样本GPT-4o。
  • 适合中文自然语言处理、媒体分析与舆论研究者使用。

新闻框架是影响公众对时事认知的重要手段。尽管已有多种语言的自动新闻框架检测数据集,但尚无聚焦中文的数据集,而中文具有复杂的字义和独特的语言特征。本文提出首个中文新闻框架数据集,可独立使用或作为SemEval-2023任务3的补充资源。我们详述数据构建过程,并开展基线实验,验证该数据集的必要性并建立未来研究基准。实验采用微调XLM-RoBERTa-Base和GPT-4o零样本设置,结果显示:在所有语言上,GPT-4o表现显著逊于微调后的XLM-RoBERTa。仅使用本数据集时,中文新闻框架检测的F1-micro得分为0.719;结合SemEval数据集后提升至0.753。正向新闻框架检测结果良好,表明该数据集对中文新闻框架识别具有重要价值,是SemEval-2023任务3的有力补充。

原文摘要 · Abstract (English)

Framing is an essential device in news reporting, allowing the writer to influence public perceptions of current affairs. While there are existing automatic news framing detection datasets in various languages, none of them focus on news framing in the Chinese language which has complex character meanings and unique linguistic features. This study introduces the first Chinese News Framing dataset, to be used as either a stand-alone dataset or a supplementary resource to the SemEval-2023 task 3 dataset. We detail its creation and we run baseline experiments to highlight the need for such a dataset and create benchmarks for future research, providing results obtained through fine-tuning XLM-RoBERTa-Base and using GPT-4o in the zero-shot setting. We find that GPT-4o performs significantly worse than fine-tuned XLM-RoBERTa across all languages. For the Chinese language, we obtain an F1-micro (the performance metric for SemEval task 3, subtask 2) score of 0.719 using only samples from our Chinese News Framing dataset and a score of 0.753 when we augment the SemEval dataset with Chinese news framing samples. With positive news frame detection results, this dataset is a valuable resource for detecting news frames in the Chinese language and is a valuable supplement to the SemEval-2023 task 3 dataset.

新闻框架中文NLP数据集舆论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。