arXiv:2601.07765cs.CL2026-01Conference of the …被引 4

用故事孪生对对比学习,自动识别叙事中关键情节。

Contrastive Learning with Narrative Twins for Modeling Story Salience

  • 构建叙事孪生对:同一剧情不同表达的故事对
  • 对比学习使模型能区分核心情节与干扰内容
  • 总结操作最可靠,适合长篇故事分析

理解叙事需识别推动故事发展的关键事件。我们提出一种对比学习框架,通过叙事孪生对——共享相同剧情但表面形式不同的故事——学习故事嵌入。模型训练目标是区分一个故事与其孪生对及具有相似表面特征但剧情不同的干扰项。利用所得嵌入,评估四种基于叙事理论的操作(删除、移位、破坏、摘要)以推断情节显著性。在ROCStories短故事和维基百科长篇剧情摘要上的实验表明,对比学习得到的嵌入优于掩码语言模型基线,且摘要操作对识别显著句最为可靠。若无现成孪生对,可通过随机丢弃生成;有效干扰项可通过提示大模型或使用同一篇长故事的不同部分获得。

原文摘要 · Abstract (English)

Understanding narratives requires identifying which events are most salient for a story's progression. We present a contrastive learning framework for modeling narrative salience that learns story embeddings from narrative twins: stories that share the same plot but differ in surface form. Our model is trained to distinguish a story from both its narrative twin and a distractor with similar surface features but different plot. Using the resulting embeddings, we evaluate four narratologically motivated operations for inferring salience (deletion, shifting, disruption, and summarization). Experiments on short narratives from the ROCStories corpus and longer Wikipedia plot summaries show that contrastively learned story embeddings outperform a masked-language-model baseline, and that summarization is the most reliable operation for identifying salient sentences. If narrative twins are not available, random dropout can be used to generate the twins from a single story. Effective distractors can be obtained either by prompting LLMs or, in long-form narratives, by using different parts of the same story.

对比学习叙事理解情节显著性孪生对

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。