用讲座上下文提升中文技术讲座中英文术语识别准确率。
Context-Aware ASR for Mandarin Technical Lectures
- 自建术语词典,通过两阶段解码增强术语识别
- 在Breeze-ASR-25上术语召回率从52.50%提至60.13%
- 新评估方式可发现传统错误率忽略的术语误识
技术讲座混合使用普通话与英文术语,这些术语虽占字符极少却承载核心语义,导致字符错误率(CER)无法反映其识别失败。本文构建了一个富含术语的中文人工智能/机器学习讲座数据集,并提出以术语为中心的评估指标。提出一种无需参考文本的两阶段解码方法:第一阶段仅分段识别,提取高频术语生成自建词典;第二阶段用该词典引导识别。在五个ASR模型上,该方法均提升术语召回率且不增加或降低CER。在Breeze-ASR-25上,术语召回率达60.13%,结合外部小词表的混合方法更达62.05%召回率与82.73%术语精度。模型自身输出恢复的讲座上下文成为有效信号,术语中心评估揭示了传统CER遗漏的错误。
原文摘要 · Abstract (English)
Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) therefore hides their recognition failures. We study whether lecture context helps recognize these terms. We build a term-rich Mandarin AI/ML lecture benchmark, and we define term-centric metrics that measure technical-term recognition directly. We then propose a two-pass, reference-free decoding method. The first pass runs segment-only ASR. We extract the most frequent technical terms from the first-pass hypotheses, and we prompt the recognizer with this self-built glossary in the second pass. Across five ASR backbones, the first-pass glossary raises term recall for every model and holds or lowers CER on all five. On Breeze-ASR-25 it lifts term recall from 52.50% to 60.13% while lowering CER, and a hybrid that adds a small external term list reaches 62.05% recall and 82.73% term precision. Lecture context, recovered from the model's own output, is a practical signal for technical-term recognition. Term-centric evaluation exposes errors that CER misses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。