构建多语言贴纸语义数据集,提升贴纸检索准确率
Small Stickers, Big Meanings: A Multilingual Sticker Semantic Understanding Dataset with a Gamified Approach
- 设计游戏化标注框架,收集高质量贴纸查询语句
- 发布含1115条英文、615条中文查询的多语言数据集
- 适合多模态理解、人机交互与跨文化研究者使用
贴纸虽小,却是跨平台广泛使用的高度凝练视觉表达形式,深受不同文化、性别和年龄群体欢迎。然而,由于人工构建高质量贴纸查询数据集成本高、主观性强,贴纸检索仍属研究空白。尽管大语言模型在通用NLP任务中表现优异,但在处理贴纸查询这种具象化、非明确且高度情境化的语义任务时表现不佳。为此,本文提出三方面解决方案:首先,设计名为Sticktionary的游戏化标注框架,用于采集多样化、高质量且语境相关的贴纸查询;其次,构建StickerQueries数据集,包含1,115条英文和615条中文查询,由超过60名贡献者在60余小时完成标注;最后,通过大量定量与定性评估验证,该方法显著提升了查询生成质量、检索准确率与语义理解能力。为支持后续研究,本文公开发布多语言数据集及两个微调后的查询生成模型。
原文摘要 · Abstract (English)
Stickers, though small, are a highly condensed form of visual expression, ubiquitous across messaging platforms and embraced by diverse cultures, genders, and age groups. Despite their popularity, sticker retrieval remains an underexplored task due to the significant human effort and subjectivity involved in constructing high-quality sticker query datasets. Although large language models (LLMs) excel at general NLP tasks, they falter when confronted with the nuanced, intangible, and highly specific nature of sticker query generation. To address this challenge, we propose a threefold solution. First, we introduce Sticktionary, a gamified annotation framework designed to gather diverse, high-quality, and contextually resonant sticker queries. Second, we present StickerQueries, a multilingual sticker query dataset containing 1,115 English and 615 Chinese queries, annotated by over 60 contributors across 60+ hours. Lastly, Through extensive quantitative and qualitative evaluation, we demonstrate that our approach significantly enhances query generation quality, retrieval accuracy, and semantic understanding in the sticker domain. To support future research, we publicly release our multilingual dataset along with two fine-tuned query generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。