arXiv:2511.07677cs.SDcs.AI2025-11

针对儿童在教室中听不清讲话的问题,提出可实时部署的语音分离模型。

Speech Separation for Hearing-Impaired Children in the Classroom

  • 采用多通道紧凑架构MIMO-TasNet,利用空间线索分离语音。
  • 仅用一半教室数据微调,效果接近全量训练,迁移高效。
  • 适合开发儿童专用助听设备,尤其对听力障碍者实用。

课堂环境对听力障碍儿童尤为困难,背景噪声、多人对话和混响会严重降低语音感知能力。相比成人,儿童声音频谱相似度更高,分离线索更弱,而现有深度学习语音分离模型大多基于成人语音在简化、低混响条件下训练,忽略了真实教室的复杂性。本文采用MIMO-TasNet这一轻量级、低延迟的多通道架构,模拟自然教室场景,包含移动的儿童-儿童及儿童-成人说话人对,在不同噪声和距离条件下测试模型表现。通过对比成人语音训练、教室数据训练及微调模型,评估数据高效适应能力。结果表明:成人训练模型在安静环境下表现良好,但在教室中性能显著下降;使用教室数据训练后分离质量大幅提升;仅用一半教室数据微调即达到相近增益,验证了高效迁移学习的有效性。加入扩散式人声干扰(diffuse babble)训练进一步增强鲁棒性,模型在未见距离下仍保持空间感知能力。研究证明,结合空间感知架构与针对性适应策略,可有效提升儿童在嘈杂教室中的语音可懂度,为未来植入式助听设备提供技术支持。

原文摘要 · Abstract (English)

Classroom environments are particularly challenging for children with hearing impairments, where background noise, multiple talkers, and reverberation degrade speech perception. These difficulties are greater for children than adults, yet most deep learning speech separation models for assistive devices are developed using adult voices in simplified, low-reverberation conditions. This overlooks both the higher spectral similarity of children's voices, which weakens separation cues, and the acoustic complexity of real classrooms. We address this gap using MIMO-TasNet, a compact, low-latency, multi-channel architecture suited for real-time deployment in bilateral hearing aids or cochlear implants. We simulated naturalistic classroom scenes with moving child-child and child-adult talker pairs under varying noise and distance conditions. Training strategies tested how well the model adapts to children's speech through spatial cues. Models trained on adult speech, classroom data, and finetuned variants were compared to assess data-efficient adaptation. Results show that adult-trained models perform well in clean scenes, but classroom-specific training greatly improves separation quality. Finetuning with only half the classroom data achieved comparable gains, confirming efficient transfer learning. Training with diffuse babble noise further enhanced robustness, and the model preserved spatial awareness while generalizing to unseen distances. These findings demonstrate that spatially aware architectures combined with targeted adaptation can improve speech accessibility for children in noisy classrooms, supporting future on-device assistive technologies.

语音分离儿童听力助听设备空间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。