评测大模型生成适合5-9岁孩子的挪威语对话,发现多数模型语言偏成人。
Evaluating LLMs on Generating Age-Appropriate Child-Like Conversations
- 用真实儿童访谈数据做盲评,对比五款大模型生成效果。
- 评估者对5岁孩子对话识别更准,9岁孩子易误判为成人语言。
- 在低资源语言中,缺乏适龄词汇库导致模型生成内容不自然。
大型语言模型(LLMs)主要基于成人对话数据训练,在生成适合特定应用的儿童对话时面临挑战。本文比较评估了五种LLM(GPT-4、RUTER-LLAMA-2-13b、GPTSW、NorMistral-7b和NorBloom-7b)生成5岁与9岁挪威儿童适龄对话的能力。通过11位教育专业人士对真实儿童访谈数据与模型生成文本进行盲评,评估其真实性和发展适宜性。结果显示,评估者间一致性良好(ICC=0.75),且对5岁儿童的年龄判断准确率高于9岁儿童。尽管GPT-4和NorBloom-7b表现相对较好,但多数模型生成的语言被感知为比目标年龄组更复杂。这凸显了在儿童专用场景,尤其是低资源语言中,缺乏全面适龄词汇资源带来的关键挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs), predominantly trained on adult conversational data, face significant challenges when generating authentic, child-like dialogue for specialized applications. We present a comparative study evaluating five different LLMs (GPT-4, RUTER-LLAMA-2-13b, GPTSW, NorMistral-7b, and NorBloom-7b) to generate age-appropriate Norwegian conversations for children aged 5 and 9 years. Through a blind evaluation by eleven education professionals using both real child interview data and LLM-generated text samples, we assessed authenticity and developmental appropriateness. Our results show that evaluators achieved strong inter-rater reliability (ICC=0.75) and demonstrated higher accuracy in age prediction for younger children (5-year-olds) compared to older children (9-year-olds). While GPT-4 and NorBloom-7b performed relatively well, most models generated language perceived as more linguistically advanced than the target age groups. These findings highlight critical data-related challenges in developing LLM systems for specialized applications involving children, particularly in low-resource languages where comprehensive age-appropriate lexical resources are scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。