arXiv:2501.18738cs.CL2025-01

测试大模型在不同语言复杂度下的表现,发现语言难度影响判断准确性。

Examining the Robustness of Large Language Models across Language Complexity

  • 用词汇、句法、语义复杂度衡量文本,对比模型性能差异
  • 低复杂度文本上模型准确率更高,高复杂度时下降明显
  • 适合教育评估领域研究者关注模型公平性与适用边界

随着大型语言模型(LLMs)的发展,越来越多的学生模型利用它们分析学生生成的文本以理解与评估学习过程。这些模型通常使用预训练的LLM将文本转化为嵌入向量,再基于嵌入训练模型来检测特定学习构念的存在与否。然而,这些模型在处理不同语言复杂度文本时的可靠性与鲁棒性如何?在学习情境中,学生可能具有不同的语言背景和写作能力,因此必须评估模型对语言复杂度变化的鲁棒性,以确保其在各类文本上表现一致。巧合的是,已有少量但有限的研究表明语言使用确实会影响LLM性能。因此,本研究考察了基于LLM的学生模型在检测数学问题解决中的自我调节学习(SRL)时的鲁棒性。具体而言,我们通过三种语言学测量方法比较了模型在高、低词汇复杂度、句法复杂度和语义复杂度文本上的表现差异。

原文摘要 · Abstract (English)

With the advancement of large language models (LLMs), an increasing number of student models have leveraged LLMs to analyze textual artifacts generated by students to understand and evaluate their learning. These student models typically employ pre-trained LLMs to vectorize text inputs into embeddings and then use the embeddings to train models to detect the presence or absence of a construct of interest. However, how reliable and robust are these models at processing language with different levels of complexity? In the context of learning where students may have different language backgrounds with various levels of writing skills, it is critical to examine the robustness of such models to ensure that these models work equally well for text with varying levels of language complexity. Coincidentally, a few (but limited) research studies show that the use of language can indeed impact the performance of LLMs. As such, in the current study, we examined the robustness of several LLM-based student models that detect student self-regulated learning (SRL) in math problem-solving. Specifically, we compared how the performance of these models vary using texts with high and low lexical, syntactic, and semantic complexity measured by three linguistic measures.

语言模型教育评估语言复杂度自适应检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。