arXiv:2605.27805cs.CLcs.AI2026-05ACL

评测大模型对3-6岁儿童偏好理解能力,构建了2.9万条儿童人格数据集。

ChildEval: When large language models meet children's personalities

论文配图:ChildEval: When large language models meet children's personalities
图 1 · 摘自论文原文
  • 构建2.9万条3-6岁儿童人格与偏好数据,支持显式/隐式表达
  • 实验表明微调可显著提升模型对儿童偏好的响应能力
  • 适合做儿童对话系统、教育AI研究者参考

尽管大语言模型(LLMs)能实现个性化聊天,但其在面向儿童的个性化表现仍缺乏系统评估。为填补这一空白,我们提出ChildEval,一个用于评估LLM在长对话中推断并遵循儿童中心偏好能力的基准。该基准包含2.9万条3-6岁儿童的合成人格档案,提供相对静态的背景信息。每条人格关联一种儿童偏好——可能与人格一致、冲突或独立,以单句显式表达或6-10轮对话隐式表达。显式与隐式偏好反映同一潜在偏好,但表达方式不同,体现偏好表达的动态性而非人格变化。基准涵盖五大类、十四小类儿童日常生活与发展主题。我们进一步设计细粒度、以儿童为中心的评估协议,系统评估开源LLM。实验结果揭示不同个性化表示对模型响应的影响,并表明在ChildEval上微调可提升儿童中心性能。代码与数据集已公开于https://github.com/ziyanluo/ChildEval。

原文摘要 · Abstract (English)

While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.

大模型儿童认知个性化对话评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。