构建多语言新闻实体角色标注数据集,揭示报道中的叙事角色模式。
Entity Framing and Role Portrayal in the News
- 基于叙事元素设计22种角色分类,分属主角、反派、无辜三类
- 涵盖5语言1378篇新闻,超5800个实体被标注角色标签
- 支持跨语言角色识别,适合媒体分析与社会偏见研究
我们提出一个新型多语言层级新闻语料库,用于标注实体在新闻中的框架与角色呈现。该数据集采用受叙事元素启发的独特分类体系,包含22种细粒度角色(或原型),嵌套于主角、反派、无辜三类主类别中。每种角色均有明确定义,涵盖如守护者、殉道者、弱者等主角类型;暴君、欺骗者、偏执者等反派类型;以及受害者、替罪羊、被剥削者等无辜类型。数据集涵盖1,378篇近期新闻文章,涉及五种语言(保加利亚语、英语、印地语、欧洲葡萄牙语、俄语),聚焦乌克兰-俄罗斯战争与气候变化两大全球性议题。超过5,800个实体提及已标注角色标签。该数据集为角色呈现研究提供宝贵资源,并对新闻分析具有广泛意义。我们描述了数据集特征与标注流程,并报告了微调的先进多语言变换器模型及基于大语言模型的层级零样本学习在文档、段落和句子层面的评估结果。
原文摘要 · Abstract (English)
We introduce a novel multilingual hierarchical corpus annotated for entity framing and role portrayal in news articles. The dataset uses a unique taxonomy inspired by storytelling elements, comprising 22 fine-grained roles, or archetypes, nested within three main categories: protagonist, antagonist, and innocent. Each archetype is carefully defined, capturing nuanced portrayals of entities such as guardian, martyr, and underdog for protagonists; tyrant, deceiver, and bigot for antagonists; and victim, scapegoat, and exploited for innocents. The dataset includes 1,378 recent news articles in five languages (Bulgarian, English, Hindi, European Portuguese, and Russian) focusing on two critical domains of global significance: the Ukraine-Russia War and Climate Change. Over 5,800 entity mentions have been annotated with role labels. This dataset serves as a valuable resource for research into role portrayal and has broader implications for news analysis. We describe the characteristics of the dataset and the annotation process, and we report evaluation results on fine-tuned state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, a paragraph, and a sentence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。