针对印地语等语言的语音识别,提出更精准的错误诊断框架
SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR
- 通过融合领域词典实现抗拼接错误的对齐,分解各类识别错误
- 在印地语、马拉雅拉姆语和卡纳达语上验证,错误类型诊断比传统WER更准
- 适合关注语音识别细粒度评估与低资源语言研究的学者
语音识别替代打字的前提是纠错成本低于人工输入——这一阈值由错误类型决定,而非错误数量:误识别一个专业术语的修正成本远高于多打一个逗号。词错误率(WER)存在双重缺陷:它将不同类型的错误合并为单一数值,且在黏着语中结构化惩罚严重,因合法音节合并导致得分虚高。本文提出SCRIBE,一个诊断性评估框架,通过注入领域词汇库并支持音节容忍对齐,将错误分解为词汇、标点、数字和领域实体四类错误率。人工验证表明,SCRIBE的结果与专家判断一致,而WER则不匹配。我们发布了SCRIBE框架、大模型清洗流程、基准数据集及针对印地语、马拉雅拉姆语和卡纳达语的开源丰富转录模型。
原文摘要 · Abstract (English)
Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not counts: fixing a misrecognized domain term costs far more than inserting a comma. Word error rate (WER) fails on two fronts: it collapses distinct error categories into a single scalar, and it structurally penalizes agglutinative languages where valid sandhi merges inflate scores. We introduce SCRIBE, a diagnostic framework offering categorical error decomposition into lexical, punctuation, numeral, and domain-entity rates via sandhi-tolerant alignment with domain vocabulary injection. Human validation confirms SCRIBE aligns with expert judgment where WER does not. We release SCRIBE, an LLM curation pipeline, benchmarks, and open-weight rich transcription models for Hindi, Malayalam, and Kannada.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。