测试大模型模仿亲子对话能力,发现其对话表面像但深层模式差。
Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction
- 用静态与交互式双方法评估大模型模仿亲子对话能力。
- 顶尖模型在词句层面接近人类,但无法还原真实互动模式。
- 适合研究儿童语言发展或人机互动的学者参考。
大语言模型可生成类人对话,但其模拟早期儿童-成人互动的能力仍不明确。本文采用静态与交互式基准测试方法,评估大模型在多大程度上捕捉亲子对话的独特特征。结果表明,如 Llama 3 与 GPT-4o 等先进模型可在词汇和语句层面近似亲子对话,但在再现儿童与照料者的话语模式方面表现不佳,过度夸大对齐性,且多样性远低于真人。本研究旨在推动面向儿童应用的大模型综合评估基准的发展。
原文摘要 · Abstract (English)
LLMs can generate human-like dialogues, yet their ability to simulate early child-adult interactions remains largely unexplored. In this paper, we examined how effectively LLMs can capture the distinctive features of child-caregiver language in interaction, using both static and interactive benchmarking methods. We found that state-of-the-art LLMs like Llama 3 and GPT-4o can approximate child-caregiver dialogues at the word and utterance level, but they struggle to reproduce the child and caregiver's discursive patterns, exaggerate alignment, and fail to reach the level of diversity shown by humans. The broader goal of this work is to initiate the development of a comprehensive benchmark for LLMs in child-oriented applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。