arXiv:2510.20850eess.AScs.CL2025-10被引 1

大模型能听懂孩子结巴语音吗?研究给出关键答案

Can large audio language models understand child stuttering speech? speech summarization, and source separation

  • 用大音频语言模型分离孩子声音并生成保留口吃特征的摘要
  • 在单人朗读任务中摘要准确率超60%,混合对话中下降至40%以下
  • 适合临床与教育场景使用,提供可复现的提示和评估脚本

儿童语音在声学、语调和语言发展上不同于成人,口吃(重复、延长、阻塞)进一步挑战自动语音识别(ASR)和下游自然语言处理。近年来的大音频语言模型(LALMs)展现出强大的跨模态音频理解能力,但其在口吃儿童语音中的表现仍缺乏研究。我们评估了多个先进LALMs在两种场景下的表现:采访(多说话人)和朗读任务(单个儿童)。任务包括(i)单通道语音分离以提取儿童声音,(ii)仅儿童摘要生成,需保留临床相关的口吃特征并避免成人语音泄露。评估结合大语言模型作为裁判、人工专家评分和BERTScore(F1),报告模型间及模型与人类的一致性以评估可靠性。研究揭示了LALMs在混合音频中生成忠实儿童摘要的适用条件与失效场景,为临床和教育部署提供实用指导。我们提供提示和评估脚本以支持复现。

原文摘要 · Abstract (English)

Child speech differs from adult speech in acoustics, prosody, and language development, and disfluencies (repetitions, prolongations, blocks) further challenge Automatic Speech Recognition (ASR) and downstream Natural Language Processing (NLP). Recent large audio-language models (LALMs) demonstrate strong cross-modal audio understanding; however, their behavior in disfluent child speech remains underexplored. We evaluate several state-of-the-art LALMs in two settings: an interview (mixed speakers) and a reading task (single child). The tasks are (i) single-channel source separation to isolate the child and (ii) child-only summarization that preserves clinically relevant disfluencies and avoids adult-speech leakage. Evaluation combines Large Language Model (LLM) as a judge, human expert ratings, and BERTScore (F1), and we report agreement between models and between models and humans to assess reliability. Our findings delineate the conditions under which LALMs produce faithful child-only summaries from mixed audio and where they fail, offering practical guidance for clinical and educational deployments. We provide prompts and evaluation scripts to support replication.

语音识别儿童语音口吃分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。