arXiv:2602.05392cs.CLcs.AI2026-02Conference of the …

用大模型评估孩子对话,看懂他们怎么展开话题和自主发言。

Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances

  • 用大模型判断家长话类型,再从扩展性和独立性两方面评分
  • 扩展性反映逻辑深度,独立性体现孩子主导对话能力
  • 比传统长度指标更贴近真实语言发展,适合研究儿童语言进步

评估亲子对话中儿童话语质量仍面临挑战,因现有指标如平均话语长度(MLU)、词汇多样性(vocd-D)和可读性指数(Flesch-Kincaid Grade Level、Gunning Fog Index)过度依赖长度,忽略对话上下文,无法捕捉推理深度、话题延续和语篇规划等质量维度。本文提出一种基于大模型的评判框架:先分类前一句成人话语类型,再从两个维度评估儿童回应——扩展性(上下文拓展与推断深度)和独立性(推动对话进展的能力)。这两个维度对应儿童语言发展的核心特征:扩展性涵盖细节丰富、复句结构及因果、对比连接词;独立性体现主动性、话题掌控力、减少对成人支架的依赖及受众意识。通过展示年龄相关模式并提升年龄预测准确率,验证了其发展有效性;通过识别语篇关系差异,证明其语义敏感性。该框架与人工判断高度一致,支持大规模评估。它将儿童话语评价从单纯计数长度,转向衡量其在具体语境中如何有意义地推进对话。

原文摘要 · Abstract (English)

Evaluating the quality of children's utterances in adult-child dialogue remains challenging due to insufficient context-sensitive metrics. Common proxies such as Mean Length of Utterance (MLU), lexical diversity (vocd-D), and readability indices (Flesch-Kincaid Grade Level, Gunning Fog Index) are dominated by length and ignore conversational context, missing aspects of response quality such as reasoning depth, topic maintenance, and discourse planning. We introduce an LLM-as-a-judge framework that first classifies the Previous Adult Utterance Type and then scores the child's response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child's contribution to advancing the discourse). These axes reflect fundamental dimensions in child language development, where Expansion captures elaboration, clause combining, and causal and contrastive connectives. Independence captures initiative, topic control, decreasing reliance on adult scaffolding through growing self-regulation, and audience design. We establish developmental validity by showing age-related patterns and demonstrate predictive value by improving age estimation over common baselines. We further confirm semantic sensitivity by detecting differences tied to discourse relations. Our metrics align with human judgments, enabling large-scale evaluation. This shifts child utterance assessment from simply measuring length to evaluating how meaningfully the child's speech contributes to and advances the conversation within its context.

儿童语言大模型评估对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。