arXiv:2603.15897cs.CLcs.AI2026-03ACL被引 1

用大模型合集提升同义词在叙事中合理性判断的准确率

COGNAC at SemEval-2026 Task 5: LLM Ensembles for Human-Level Word Sense Plausibility Rating in Challenging Narratives

  • 结合三种提示策略,集成多个商用大模型进行推理
  • 最佳表现达0.88准确率与0.83斯皮尔曼相关系数
  • 对比提示与模型集成显著提升对人类判断的一致性

我们介绍参加SemEval-2026任务5的系统,该任务要求在5点李克特量表上评估短篇故事中同义词词义的合理性。系统通过未加权平均准确率(在人类均值一个标准差内)和斯皮尔曼等级相关系数进行评估。我们探索了三种使用多个闭源商业大模型的提示策略:(i) 基线零样本设置,(ii) 结构化推理的思维链(CoT)提示,(iii) 同时评估候选词义的对比提示。为应对标注中显著的标注者差异,我们提出通过平均模型预测结果的集成方案。最佳官方系统由三种提示策略的大模型集成构成,在排行榜中位列第4,准确率为0.88,斯皮尔曼相关系数为0.83(平均0.86)。赛后的额外实验进一步将性能提升至0.92准确率与0.85斯皮尔曼相关系数(平均0.89)。研究发现,对比提示在不同模型族中均稳定提升表现,模型集成显著增强与人类均值判断的一致性,表明大模型集成特别适用于涉及多标注者的主观语义评估任务。

原文摘要 · Abstract (English)

We describe our system for SemEval-2026 Task 5, which requires rating the plausibility of given word senses of homonyms in short stories on a 5-point Likert scale. Systems are evaluated by the unweighted average of accuracy (within one standard deviation of mean human judgments) and Spearman Rank Correlation. We explore three prompting strategies using multiple closed-source commercial LLMs: (i) a baseline zero-shot setup, (ii) Chain-of-Thought (CoT) style prompting with structured reasoning, and (iii) a comparative prompting strategy for evaluating candidate word senses simultaneously. Furthermore, to account for the substantial inter-annotator variation present in the gold labels, we propose an ensemble setup by averaging model predictions. Our best official system, comprising an ensemble of LLMs across all three prompting strategies, placed 4th on the competition leaderboard with 0.88 accuracy and 0.83 Spearman's rho (0.86 average). Post-competition experiments with additional models further improved this performance to 0.92 accuracy and 0.85 Spearman's rho (0.89 average). We find that comparative prompting consistently improved performance across model families, and model ensembling significantly enhanced alignment with mean human judgments, suggesting that LLM ensembles are especially well suited for subjective semantic evaluation tasks involving multiple annotators.

大模型集成语义评估同义词判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。