提出新评估方法,精准定位语音去口误模型的失效环节。
Z-Scores: A Metric for Linguistically Assessing Disfluency Removal
- 基于语义片段的分类评估,区分编辑、插入、插入语三类口误。
- 揭示大模型在插入语和插入型口误上表现差,传统指标掩盖此问题。
- 帮助研究者针对性优化提示词或数据,提升模型鲁棒性。
语音去口误评估仅靠整体词级指标(如精确率、召回率、F1)不够。传统词汇级指标虽能反映总体性能,却无法揭示模型成功或失败的原因。本文提出Z-Scores,一种基于语言学的片段级评估指标,可对不同口误类型(EDITED、INTJ、PRN)进行行为分析。其确定性对齐模块实现生成文本与有口误原始语句间的稳健映射,使Z-Scores能够暴露词级指标所隐藏的系统性缺陷。通过提供分类型诊断,Z-Scores助力研究者识别模型失效模式,并设计针对性干预措施(如定制提示或数据增强),带来可量化的性能提升。以大模型为例的案例研究显示,Z-Scores揭示了聚合F1值掩盖的INTJ和PRN类口误处理难题,直接指导模型优化策略。
原文摘要 · Abstract (English)
Evaluating disfluency removal in speech requires more than aggregate token-level scores. Traditional word-based metrics such as precision, recall, and F1 (E-Scores) capture overall performance but cannot reveal why models succeed or fail. We introduce Z-Scores, a span-level linguistically-grounded evaluation metric that categorizes system behavior across distinct disfluency types (EDITED, INTJ, PRN). Our deterministic alignment module enables robust mapping between generated text and disfluent transcripts, allowing Z-Scores to expose systematic weaknesses that word-level metrics obscure. By providing category-specific diagnostics, Z-Scores enable researchers to identify model failure modes and design targeted interventions -- such as tailored prompts or data augmentation -- yielding measurable performance improvements. A case study with LLMs shows that Z-Scores uncover challenges with INTJ and PRN disfluencies hidden in aggregate F1, directly informing model refinement strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。