评测顶尖大模型在时态、否定等语言细节上的理解能力,发现Compound-Beta表现最佳。
Does Language Model Understand Language?
- 构建双语语料库LUCID,针对时态、否定等语言现象设计挑战性句子对。
- Compound-Beta在英语和跨语言场景中均达最高皮尔逊相关系数,误差最小。
- 提出新指标HCE准确率,衡量模型预测与人类判断的接近程度。
尽管自然语言生成与理解取得进展,大语言模型(LM)在时态、否定、语态和情态等细微语言现象上仍表现不佳,这些正是有效人类沟通的核心。在联合国可持续发展目标4(教育公平)背景下,教育技术中部署大模型需严格评估其语言理解能力。随着大模型广泛用于辅导系统、自动评分和翻译,其与人类语言理解的一致性至关重要。本研究评估了包括MISTRAL-SABA-24B、LLaMA-4-Scout-17B、LLaMA-3.3-70B、Gemma2-9B和Compound-Beta在内的主流模型,在英语与孟加拉语中的表现。我们引入系统化评估框架,并构建了名为LUCID的新数据集,专门测试模型在否定、时态、语态等方面的推理能力。采用皮尔逊相关、斯皮尔曼相关、平均绝对误差及新颖的基于语言学的HCE准确率进行评估。HCE准确率衡量模型预测是否落在人类评分均值的一个标准差内,反映对语言解释变异性的容忍度。结果表明,Compound-Beta表现最均衡,英文下皮尔逊相关最高,在多语言数据上也表现出稳健性能,显示出强跨语言一致性。
原文摘要 · Abstract (English)
Despite advances in natural language generation and understanding, LM still struggle with fine grained linguistic phenomena such as tense, negation, voice, and modality which are the elements central to effective human communication. In the context of the United Nations SDG 4, where linguistic clarity is critical, the deployment of LMs in educational technologies demands careful scrutiny. As LMs are increasingly powering applications like tutoring systems, automated grading, and translation, their alignment with human linguistic interpretation becomes essential for effective learning. In this study, we conduct a evaluation of SOTA language models across these challenging contexts in both English and Bengali. To ensure a structured assessment, we introduce a new Route for Evaluation of Cognitive Inference in Systematic Environments guidelines. Our proposed LUCID dataset, composed of carefully crafted sentence pairs in English and Bengali, specifically challenges these models on critical aspects of language comprehension, including negation, tense, voice variations. We assess the performance of SOTA models including MISTRAL-SABA-24B, LLaMA-4-Scout-17B, LLaMA-3.3-70B, Gemma2-9B, and Compound-Beta using standard metrics like Pearson correlation, Spearman correlation, and Mean Absolute Error, as well as novel, linguistically inspired metric the HCE accuracy. The HCE accuracy measures how often model predictions fall within one standard deviation of the mean human rating, thus capturing human like tolerance for variability in language interpretation. Our findings highlight Compound-Beta as the most balanced model, consistently achieving high correlations and low MAEs across diverse language conditions. It records the highest Pearson correlation in English and demonstrates robust performance on mixed-language data, indicating a strong alignment with human judgments in cross lingual scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。