构建故事相似性评估任务,推动叙事表征学习发展
SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning
- 将故事相似性定义为二分类问题,基于多人标注验证
- 收集超1000组故事三元组,每组至少两人为一致标注
- 大模型集成在分类任务中表现领先,嵌入模型优化效果接近微调
我们提出了叙事相似性与叙事表征学习的共享任务NSNRL(读作'nass-na-rel')。该任务将叙事相似性操作化为二分类问题:判断两个故事中哪一个更接近锚点故事。我们提出了一种兼容叙事理论与直觉判断的新定义。基于该定义收集了超过1000个故事摘要三元组的标注数据,每个三元组至少有两个标注者达成一致。本文描述了数据采样与标注流程,并对提交的系统进行了综述。共收到来自46支队伍的71个最终提交,覆盖两个赛道。在三元组分类设置中,大型语言模型集成系统表现优异;在嵌入表示设置中,预训练嵌入模型经预处理和后处理后的性能与自定义微调方案相当。分析表明两类系统仍有提升空间。任务官网提供了所有团队的嵌入可视化及实例级分类结果。
原文摘要 · Abstract (English)
We present the shared task on narrative similarity and narrative representation learning - NSNRL (pronounced "nass-na-rel"). The task operationalizes narrative similarity as a binary classification problem: determining which of two stories is more similar to an anchor story. We introduce a novel definition of narrative similarity, compatible with both narrative theory and intuitive judgment. Based on the similarity judgments collected under this concept, we also evaluate narrative embedding representations. We collected at least two annotations each for more than 1,000 story summary triples, with each annotation being backed by at least two annotators in agreement. This paper describes the sampling and annotation process for the dataset; further, we give an overview of the submitted systems and the techniques they employ. We received a total of 71 final submissions from 46 teams across our two tracks. In our triple-based classification setup, LLM ensembles make up many of the top-scoring systems, while in the embedding setup, systems with pre- and post-processing on pretrained embedding models perform about on par with custom fine-tuned solutions. Our analysis identifies potential headroom for improvement of automated systems in both tracks. The task website includes visualizations of embeddings alongside instance-level classification results for all teams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。