用结构化提示分解故事组件,提升词语义合理性评分准确率
NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating
- 将故事拆解为前情、目标句、结尾三部分,分步评估语义合理性
- 结构化提示+明确判断规则使性能超越微调模型和嵌入方法
- 提示设计比模型规模更重要,适合需要可解释性的任务
词语义合理性评分要求在包含歧义同音词的简短叙事故事中,预测人类感知的语义合理性(1-5分)。本文系统比较三种方法:(1) 嵌入方法结合句子嵌入与标准回归器,(2) 使用参数高效适配的Transformer微调,(3) 大语言模型(LLM)提示结合结构化推理与显式决策规则。最优系统采用结构化提示策略,将评估分解为叙事成分(前情、目标句、结尾),并应用显式决策规则进行评分校准。分析表明,结构化提示配合决策规则优于微调模型和嵌入方法,且提示设计的重要性超过模型规模。
原文摘要 · Abstract (English)
Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short narrative stories containing ambiguous homonyms. This paper systematically compares three approaches: (1) embedding-based methods pairing sentence embeddings with standard regressors, (2) transformer fine-tuning with parameter-efficient adaptation, and (3) large language model (LLM) prompting with structured reasoning and explicit decision rules. The best-performing system employs a structured prompting strategy that decomposes evaluation into narrative components (precontext, target sentence, ending) and applies explicit decision rules for rating calibration. The analysis reveals that structured prompting with decision rules outperforms both fine-tuned models and embedding-based approaches, and that prompt design matters more than model scale for this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。