构建首个中文隐喻情绪细粒度标注数据集,助力跨模态情感理解
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
- 采集5000个中英文图文隐喻广告对,标注隐喻与情绪关系
- 涵盖10类细粒度情绪(如喜悦、愤怒、惊讶等),支持精准分类
- 填补中文多模态隐喻情绪研究空白,适合情感计算与NLP研究者
隐喻在表达情感中起关键作用,是情感智能的重要组成部分。随着多模态数据的普及和传播方式的多样化,多模态隐喻大量涌现,使情绪分类比单模态场景更复杂。然而,针对多模态隐喻细粒度情绪数据集的构建研究匮乏,且现有工作主要集中在英文语境,忽视了语言间情感细微差异。为弥补这一空白,我们构建了一个中文多模态数据集EmoMeta,包含5000个文本-图像配对的隐喻广告样本。每条数据均经过精细标注,涵盖隐喻存在性、领域关系及10类细粒度情绪:喜悦、爱、信任、恐惧、悲伤、厌恶、愤怒、惊讶、期待和中性。该数据集已开源(https://github.com/DUTIR-YSQ/EmoMeta),可推动该新兴领域的持续发展。
原文摘要 · Abstract (English)
Metaphors play a pivotal role in expressing emotions, making them crucial for emotional intelligence. The advent of multimodal data and widespread communication has led to a proliferation of multimodal metaphors, amplifying the complexity of emotion classification compared to single-mode scenarios. However, the scarcity of research on constructing multimodal metaphorical fine-grained emotion datasets hampers progress in this domain. Moreover, existing studies predominantly focus on English, overlooking potential variations in emotional nuances across languages. To address these gaps, we introduce a multimodal dataset in Chinese comprising 5,000 text-image pairs of metaphorical advertisements. Each entry is meticulously annotated for metaphor occurrence, domain relations and fine-grained emotion classification encompassing joy, love, trust, fear, sadness, disgust, anger, surprise, anticipation, and neutral. Our dataset is publicly accessible (https://github.com/DUTIR-YSQ/EmoMeta), facilitating further advancements in this burgeoning field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。