研究形容词修饰如何改变句子合理性,揭示模型理解微妙语义的短板。
Modelling Adjectival Modification Effects on Semantic Plausibility
- 用句向量模型分析形容词修改对语义合理性的影响
- 16000组句子对比显示句向量模型表现不如RoBERTa
- 强调需更均衡的评估方法以提升结果可信度
尽管事件合理性判断任务如“新闻相关”已有较多研究,但关于事件因修饰词改变而引发的合理性变化关注较少。理解此类变化对对话生成、常识推理和幻觉检测等任务至关重要,例如可区分朋友间‘温和讽刺’是亲昵而非无礼。本文针对包含16,000组仅差一个形容词修饰的英语句子对的ADEPT挑战基准,提出一种基于句向量的新建模方法。实验表明,尽管句向量在概念上契合该任务,但其表现反而逊于RoBERTa等基于Transformer的模型。深入对比前人工作进一步指出:现有评估存在数据分布不均问题,导致模型性能误判,削弱了结果可信度。
原文摘要 · Abstract (English)
While the task of assessing the plausibility of events such as ''news is relevant'' has been addressed by a growing body of work, less attention has been paid to capturing changes in plausibility as triggered by event modification. Understanding changes in plausibility is relevant for tasks such as dialogue generation, commonsense reasoning, and hallucination detection as it allows to correctly model, for example, ''gentle sarcasm'' as a sign of closeness rather than unkindness among friends [9]. In this work, we tackle the ADEPT challenge benchmark [6] consisting of 16K English sentence pairs differing by exactly one adjectival modifier. Our modeling experiments provide a conceptually novel method by using sentence transformers, and reveal that both they and transformer-based models struggle with the task at hand, and sentence transformers - despite their conceptual alignment with the task - even under-perform in comparison to models like RoBERTa. Furthermore, an in-depth comparison with prior work highlights the importance of a more realistic, balanced evaluation method: imbalances distort model performance and evaluation metrics, and weaken result trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。