arXiv:2412.06179cs.CLcs.AI2024-12被引 2

为食物类推文构建多维度标注数据集,助力自然语言理解研究

Annotations for Exploring Food Tweets From Multiple Aspects

  • 基于12年积累的300万条食物推文,新增多任务人工标注数据
  • 涵盖机器翻译、实体识别等4个任务,均提供时间平衡的情感分析数据
  • 适合做多模态、跨任务自然语言处理研究的学者和开发者

本研究基于持续收集超过12年的拉脱维亚推特饮食语料库(LTEC),该语料库聚焦于食物、饮品、进食与饮水相关的推文,目前已包含近300万条带有基础信息及自动与人工标注元数据的推文。本文在LTEC基础上,补充了针对机器翻译、命名实体识别、时间平衡情感分析以及文本-图像关系分类的多任务人工标注评估数据集。我们使用基线模型对各数据集进行实验,并指出了各类建模方法未来面临的挑战。

原文摘要 · Abstract (English)

This research builds upon the Latvian Twitter Eater Corpus (LTEC), which is focused on the narrow domain of tweets related to food, drinks, eating and drinking. LTEC has been collected for more than 12 years and reaching almost 3 million tweets with the basic information as well as extended automatically and manually annotated metadata. In this paper we supplement the LTEC with manually annotated subsets of evaluation data for machine translation, named entity recognition, timeline-balanced sentiment analysis, and text-image relation classification. We experiment with each of the data sets using baseline models and highlight future challenges for various modelling approaches.

推文分析多任务学习数据标注情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。