探索语音大模型对儿童口吃语音的语义理解能力,发现其推理性能在复杂场景下显著下降。
Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech
- 使用指令引导模型聚焦儿童说话者并保留口吃特征
- 在混合说话人访谈中,模型对口吃语音的语义推理准确率下降明显
- 适用于临床语音分析与儿童言语障碍研究的模型评估
儿童语音在声学、韵律和语言结构上不同于成人语音,言语不流畅(如重复)进一步挑战自动理解。尽管语音语言模型(ALMs)展现出从语音中进行强语义推理的能力,但其在混合说话人场景下对口吃儿童语音的推理能力仍未知。我们通过两项任务进行探究:以儿童为中心的语义摘要与语音蕴含判断。实验基于未显式区分说话人的儿童口吃者混合对话录音。模型通过指令引导聚焦儿童说话者,保留临床相关不流畅性,并避免成人语音干扰。评估结合大语言模型评判与基于参考文本的指标,以转录-真值基线为锚点分离错误。结果表明,尽管ALMs能从口吃语音中提取高层语义,但随着不流畅程度增加,推理性能显著下降。
原文摘要 · Abstract (English)
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。