用大模型分析儿童科学画作,建立标准化参考体系
Constructing a Norm for Children's Scientific Drawing: Distribution Features Based on Semantic Similarity of Large Language Models
- 用大模型与词向量计算画作语义相似度,发现多数画作相似度超0.8
- 画作一致性独立于识别准确率,存在一致性偏差现象
- 教学实验内容影响儿童绘图重点,更关注实验过程而非概念解释
利用儿童画作评估其科学概念理解已被证明有效,但以往研究存在两大问题:一是画作内容高度依赖任务设计,结论生态效度低;二是画作解读过度依赖研究者主观判断。为此,本研究使用大语言模型(LLM)分析1420幅覆盖9个科学主题的儿童科学画作,并通过word2vec算法计算其语义相似度,探索儿童对同一主题是否具有稳定的表现形式,尝试建立儿童科学画作的标准参照体系。结果表明,多数画作表现具有一致性,语义相似度大多高于0.8;且一致性不受LLM识别准确率影响,揭示了一致性偏差的存在。后续分析中,采用肯德尔等级相关系数探讨样本量、抽象程度和焦点点对画作的影响,并通过词频统计考察儿童是否再现课堂所授内容。发现识别准确率是最敏感指标,样本量与语义相似度等数据均与其相关;同时,课堂实验与教学目标的一致性也是重要因素,许多学生更关注实验本身而非其所解释的概念。
原文摘要 · Abstract (English)
The use of children's drawings to examining their conceptual understanding has been proven to be an effective method, but there are two major problems with previous research: 1. The content of the drawings heavily relies on the task, and the ecological validity of the conclusions is low; 2. The interpretation of drawings relies too much on the subjective feelings of the researchers. To address this issue, this study uses the Large Language Model (LLM) to identify 1420 children's scientific drawings (covering 9 scientific themes/concepts), and uses the word2vec algorithm to calculate their semantic similarity. The study explores whether there are consistent drawing representations for children on the same theme, and attempts to establish a norm for children's scientific drawings, providing a baseline reference for follow-up children's drawing research. The results show that the representation of most drawings has consistency, manifested as most semantic similarity>0.8. At the same time, it was found that the consistency of the representation is independent of the accuracy (of LLM's recognition), indicating the existence of consistency bias. In the subsequent exploration of influencing factors, we used Kendall rank correlation coefficient to investigate the effects of "sample size", "abstract degree", and "focus points" on drawings, and used word frequency statistics to explore whether children represented abstract themes/concepts by reproducing what was taught in class. It was found that accuracy (of LLM's recognition) is the most sensitive indicator, and data such as sample size and semantic similarity are related to it; The consistency between classroom experiments and teaching purpose is also an important factor, many students focus more on the experiments themselves rather than what they explain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。