arXiv:2505.17950cs.CLcs.AI2025-05被引 2

测试主流模型对学生科学文本中符号表达的处理能力,发现GPT嵌入模型表现最佳。

Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts

  • 用物理题学生作答中的公式表达测试不同嵌入模型
  • GPT-text-embedding-3-large在相似度和下游任务中均领先
  • 提醒教育数据研究者选模型时需考虑符号表达处理能力

近年来,自然语言处理(NLP)已成为教育数据挖掘的关键工具,尤其用于分析学生生成的语言内容。为支持量化研究与评估,通常采用嵌入模型将文本转化为数值表示,以捕捉语义信息。然而,在科学类语言中,方程、公式等符号表达带来挑战,现有嵌入模型普遍难以有效处理。当前研究与应用常忽视此问题或直接移除符号,可能导致研究偏差与应用性能下降。本研究系统评估了多种现代嵌入模型在处理物理相关符号表达上的能力,基于真实学生作答中的科学符号数据,通过两种方式评价:1)基于相似性的分析;2)集成至机器学习流程。结果表明各模型表现差异显著,其中OpenAI的GPT-text-embedding-3-large优于所有其他模型,但优势程度适中而非压倒性。研究强调,教育数据挖掘领域研究人员与实践者在处理含符号表达的科学语言时,必须谨慎选择嵌入模型。代码与(部分)数据已公开于https://doi.org/10.17605/OSF.IO/6XQVG。

原文摘要 · Abstract (English)

In recent years, natural language processing (NLP) has become integral to educational data mining, particularly in the analysis of student-generated language products. For research and assessment purposes, so-called embedding models are typically employed to generate numeric representations of text that capture its semantic content for use in subsequent quantitative analyses. Yet when it comes to science-related language, symbolic expressions such as equations and formulas introduce challenges that current embedding models struggle to address. Existing research studies and practical applications often either overlook these challenges or remove symbolic expressions altogether, potentially leading to biased research findings and diminished performance of practical applications. This study therefore explores how contemporary embedding models differ in their capability to process and interpret science-related symbolic expressions. To this end, various embedding models are evaluated using physics-specific symbolic expressions drawn from authentic student responses, with performance assessed via two approaches: 1) similarity-based analyses and 2) integration into a machine learning pipeline. Our findings reveal significant differences in model performance, with OpenAI's GPT-text-embedding-3-large outperforming all other examined models, though its advantage over other models was moderate rather than decisive. Overall, this study underscores the importance for educational data mining researchers and practitioners of carefully selecting NLP embedding models when working with science-related language products that include symbolic expressions. The code and (partial) data are available at https://doi.org/10.17605/OSF.IO/6XQVG.

NLP嵌入教育数据符号表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。