arXiv:2606.10454eess.AScs.SD2026-06中稿 · Interspeech 2026

首个统一识别儿童与成人语音的语音大模型框架,提升跨年龄识别效果。

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

论文配图:Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR
图 1 · 摘自论文原文
  • 采用专家混合架构,通过分类器路由动态分配不同年龄语音处理任务。
  • 在多个儿童语音数据集上超越基线,同时保持成人语音识别性能。
  • 引入熵感知路由机制,解决年龄边界处识别不确定性问题,适合多场景语音系统部署。

尽管语音大模型在成人语音识别方面表现优异,但其在儿童语音上的应用仍不充分,单一模型难以同时处理多样化的成人与儿童语音。本文提出一种基于专家混合(MoE)的语音大模型框架,实现跨年龄、跨环境的统一语音识别。该框架采用基于分类器的领域路由(C-DR),结合粗粒度到细粒度的策略,并集成混合投影器(MoP)与混合低秩适配(MoL)以建模不同领域差异。为缓解领域边界处的路由不确定性,引入熵感知路由(EAR)机制,动态融合共享专家。在公开儿童语音数据集上的实验表明,该方法持续优于基线,且不影响成人语音识别性能。据我们所知,这是首个利用语音大模型实现涵盖儿童与成人的统一多领域语音识别的工作。

原文摘要 · Abstract (English)

While Speech Large Language Models (Speech-LLMs) have achieved strong performance on adult Automatic Speech Recognition (ASR), their effectiveness on child speech remains under-explored, and single models often struggle to handle diverse adult and child age groups simultaneously. This paper proposes a Mixture-of-Experts (MoE) Speech-LLM for unified ASR across adult and child speech spanning diverse environments and age groups. The framework employs a Classifier-based Domain Router (C-DR) with a coarse-to-fine strategy and integrates both a Mixture-of-Projectors (MoP) and a Mixture-of-LoRAs (MoL) to model domain-specific variations. To address routing uncertainty near domain boundaries, an Entropy-Aware Routing (EAR) mechanism is introduced to dynamically incorporate a shared expert. Experiments on public child corpora demonstrate consistent improvements over baselines while preserving adult ASR performance. To our knowledge, this is the first work leveraging Speech-LLMs for unified, multi-domain ASR encompassing both children and adults.

语音识别专家混合儿童语音大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。