构建首个俄语新闻标题谐音数据集,助力大模型理解隐喻语言。
KoWit-24: A Richly Annotated Dataset of Wordplay in News Headlines
- 标注2700条俄语新闻标题的双关语类型与指向
- 发现构式转换是主流双关机制,此前数据集未充分覆盖
- 适合研究幽默理解、多模态推理的NLP学者使用
我们提出KoWit-24,一个包含2700条俄语新闻标题的细粒度双关语标注数据集。该数据集标注了双关语的存在性、类型、锚点词及所指词语或短语。与多数现成笑话数据集不同,KoWit-24提供上下文——每条标题均配有新闻导语和摘要。数据集中最常见的双关类型为习语、搭配结构和专有名词的转化,这一机制在以往幽默数据集中被严重低估。五种大语言模型的实验表明,双关语检测与理解任务仍有巨大提升空间。数据集与评估脚本已开源:https://github.com/Humor-Research/KoWit-24。
原文摘要 · Abstract (English)
We present KoWit-24, a dataset with fine-grained annotation of wordplay in 2,700 Russian news headlines. KoWit-24 annotations include the presence of wordplay, its type, wordplay anchors, and words/phrases the wordplay refers to. Unlike the majority of existing humor collections of canned jokes, KoWit-24 provides wordplay contexts -- each headline is accompanied by the news lead and summary. The most common type of wordplay in the dataset is the transformation of collocations, idioms, and named entities -- the mechanism that has been underrepresented in previous humor datasets. Our experiments with five LLMs show that there is ample room for improvement in wordplay detection and interpretation tasks. The dataset and evaluation scripts are available at https://github.com/Humor-Research/KoWit-24
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。