arXiv:2603.22709cs.CLeess.AS2026-03

对比大模型与模块化系统在多说话人语音识别中的表现

Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics

  • 用语义相似度替代编辑距离,改进重叠语音评估
  • 双说话人下大模型表现尚可,但人数和重叠增加时性能下降
  • 模块化系统在复杂场景下更稳定,适合实际对话应用

对话式自动语音识别因重叠语音、远场噪声和说话人数量变化而面临挑战。尽管基于大语言模型(LLM)的系统在单说话人基准上表现良好,但在多说话人场景下的鲁棒性尚不明确。本文从重叠鲁棒性、语义保真度、说话人数量及单/多通道输入四个维度,系统比较了基于大模型与模块化流水线方法。为捕捉传统指标忽略的语义错误,提出tcpSemER,将tcpWER中的Levenshtein距离替换为基于嵌入的语义相似度,并进一步分解为重叠与非重叠成分以实现细粒度分析。在三个数据集上的实验表明,大模型系统在双说话人场景下具有竞争力,但随着说话人数量和重叠程度增加,性能显著下降;而模块化流水线则保持更高鲁棒性。

原文摘要 · Abstract (English)

Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well on single-speaker benchmarks, their robustness in multi-speaker settings is unclear. We systematically compare LLM-based and modular pipeline approaches along four axes: overlap robustness, semantic fidelity, speaker count, and single- versus multi-channel input. To capture meaning-altering errors that conventional metrics miss, we introduce tcpSemER, which extends tcpWER by replacing Levenshtein distance with embedding-based semantic similarity. We further decompose tcpWER into overlapping and non-overlapping components for finer-grained analysis. Experiments across three datasets show that LLM-based systems are competitive in two-speaker settings but degrade as speaker count and overlap increase, whereas modular pipelines remain more robust.

语音识别多说话人语义评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。