大模型在理解抽象语义上表现不佳,新方法提升准确率超3%。
LLMs Struggle with Abstract Meaning Comprehension More Than Expected
- 用双向注意力机制模拟人类认知,动态关注文本与选项
- 在ReCAM任务中,模型准确率提升4.06%和3.41%
- 适合研究抽象语义、模型可解释性的学者参考
理解抽象意义对高级语言理解至关重要。尽管研究广泛,抽象词汇因其非具体、高阶语义仍具挑战性。SemEval-2021 Task 4(ReCAM)通过填空式题目评估模型对抽象概念的理解能力,包含篇章与五个抽象选项。关键发现:(1) 多数大语言模型(包括GPT-4o)在零样本、单样本及少样本设置下均表现不佳,而微调模型如BERT和RoBERTa表现更优。(2) 提出一种受人类认知策略启发的双向注意力分类器,通过动态关注篇章与选项,使微调模型在任务1上准确率提升4.06%,任务2上提升3.41%,展现其在抽象语义理解中的潜力。
原文摘要 · Abstract (English)
Understanding abstract meanings is crucial for advanced language comprehension. Despite extensive research, abstract words remain challenging due to their non-concrete, high-level semantics. SemEval-2021 Task 4 (ReCAM) evaluates models' ability to interpret abstract concepts by presenting passages with questions and five abstract options in a cloze-style format. Key findings include: (1) Most large language models (LLMs), including GPT-4o, struggle with abstract meaning comprehension under zero-shot, one-shot, and few-shot settings, while fine-tuned models like BERT and RoBERTa perform better. (2) A proposed bidirectional attention classifier, inspired by human cognitive strategies, enhances fine-tuned models by dynamically attending to passages and options. This approach improves accuracy by 4.06 percent on Task 1 and 3.41 percent on Task 2, demonstrating its potential for abstract meaning comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。