arXiv:2604.16262cs.CL2026-04ACL

用大模型评估故事中词语意义的合理性,贴近人类判断。

SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation

论文配图:SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation
图 1 · 摘自论文原文
  • 基于结构化推理框架,利用大模型分析故事语境下的词义合理性。
  • 动态少样本提示使商用大模型得分接近人类评分,准确率超基准方法。
  • 集成多个模型可更好模拟多人一致性,适合高精度语义理解任务。

近年来,语言模型在自然语言理解方面取得显著进展。尽管广泛使用的基准测试表明大型语言模型(LLMs)能有效进行词义消歧,但其在真实叙事场景中的实际应用仍待深入探索。SemEval-2026 Task 5 通过预测短篇故事中词义的人类感知合理性,填补了这一空白。本文提出一种基于大模型的框架,用于叙事文本中同义词义的合理性评分,采用结构化推理机制。我们考察了不同推理策略下微调低参数模型的效果,以及大参数模型使用动态少样本提示的影响,以实现精准的词义识别与合理性估计。实验结果表明,使用动态少样本提示的商用大参数模型能紧密复现人类的合理性判断。此外,模型集成略微提升性能,更准确模拟五位人工标注者的一致性模式。

原文摘要 · Abstract (English)

Recent advances in language models have substantially improved Natural Language Understanding (NLU). Although widely used benchmarks suggest that Large Language Models (LLMs) can effectively disambiguate, their practical applicability in real-world narrative contexts remains underexplored. SemEval-2026 Task 5 addresses this gap by introducing a task that predicts the human-perceived plausibility of a word sense within a short story. In this work, we propose an LLM-based framework for plausibility scoring of homonymous word senses in narrative texts using a structured reasoning mechanism. We examine the impact of fine-tuning low-parameter LLMs with diverse reasoning strategies, alongside dynamic few-shot prompting for large-parameter models, on accurate sense identification and plausibility estimation. Our results show that commercial large-parameter LLMs with dynamic few-shot prompting closely replicate human-like plausibility judgments. Furthermore, model ensembling slightly improves performance, better simulating the agreement patterns of five human annotators compared to single-model predictions

大模型词义消歧合理性评分叙事理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。