通过语义变化度量常识合理性,提升模型判断精度
Estimating Commonsense Plausibility through Semantic Shifts
- 用常识信息增强句子,看语义变化大小来判断合理性
- 在多个任务上优于现有方法,尤其在视觉常识任务中表现突出
- 适合需要精细常识判断的AI评估场景,如大模型评测
常识合理性评估对语言模型(LMs)至关重要,但现有生成式方法依赖似然或口头判断,在细粒度区分上表现不佳。本文提出ComPaSS,一种新型判别框架,通过测量句子添加常识信息后语义的变化程度来量化合理性:合理增补导致语义微变,不合理则引发显著偏离。在不同主干模型(包括大语言模型和视觉-语言模型)上的两类细粒度常识合理性评估任务中,ComPaSS均持续优于基线。实验表明:(1) 将ComPaSS与视觉-语言模型结合,在视觉基础常识任务中性能优于仅用语言模型;(2) 对比学习预训练能增强主干模型捕捉语义细微差别的能力,从而进一步提升ComPaSS效果。
原文摘要 · Abstract (English)
Commonsense plausibility estimation is critical for evaluating language models (LMs), yet existing generative approaches--reliant on likelihoods or verbalized judgments--struggle with fine-grained discrimination. In this paper, we propose ComPaSS, a novel discriminative framework that quantifies commonsense plausibility by measuring semantic shifts when augmenting sentences with commonsense-related information. Plausible augmentations induce minimal shifts in semantics, while implausible ones result in substantial deviations. Evaluations on two types of fine-grained commonsense plausibility estimation tasks across different backbones, including LLMs and vision-language models (VLMs), show that ComPaSS consistently outperforms baselines. It demonstrates the advantage of discriminative approaches over generative methods in fine-grained commonsense plausibility evaluation. Experiments also show that (1) VLMs yield superior performance to LMs, when integrated with ComPaSS, on vision-grounded commonsense tasks. (2) contrastive pre-training sharpens backbone models' ability to capture semantic nuances, thereby further enhancing ComPaSS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。