构建首个罗马尼亚语多模态迷因数据集,助力AI理解网络迷因。
RoMemes: A multimodal meme corpus for the Romanian language
- 构建罗马尼亚语迷因数据集,含多级标注。
- 基线模型表现有限,表明需改进AI处理能力。
- 适合研究多模态语言与社交媒体内容的学者。
迷因在在线媒体中日益流行,尤其在社交网络中广泛传播。它们通常结合图像、绘画、动画或视频与文字,传递强烈信息。为有效提取、处理和理解这些信息,人工智能应用需采用多模态算法。本文提出一个经过精心筛选的罗马尼亚语真实迷因数据集,包含多层级标注。通过基线算法验证了该数据集的可用性。结果表明,当前AI工具在处理网络迷因时仍需进一步研究和改进。
原文摘要 · Abstract (English)
Memes are becoming increasingly more popular in online media, especially in social networks. They usually combine graphical representations (images, drawings, animations or video) with text to convey powerful messages. In order to extract, process and understand the messages, AI applications need to employ multimodal algorithms. In this paper, we introduce a curated dataset of real memes in the Romanian language, with multiple annotation levels. Baseline algorithms were employed to demonstrate the usability of the dataset. Results indicate that further research is needed to improve the processing capabilities of AI tools when faced with Internet memes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。