用大模型修正失语症语音识别结果,更注重语义准确而非字面错误率。
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
- 基于大模型的裁判-编辑机制,从多个识别候选中筛选并重写不确定片段。
- 在挑战性样本上实现WER降低14.51%,语义指标提升7.66个百分点。
- 适用于需要高语义保真度的医疗语音应用,尤其适合失语症患者场景。
尽管自动语音识别(ASR)通常以词错误率(WER)为评估标准,但实际应用最终取决于语义保真度。这一差距在失语症语音中尤为严重,因发音不精准和不连贯导致严重语义扭曲。为此,我们提出一种基于大语言模型(LLM)的后处理校正代理:对top-k ASR候选进行裁判-编辑,保留高置信度片段,重写不确定部分,支持零样本与微调模式。同时,我们发布了SAP-Hypo5,目前最大的失语症语音校正基准数据集,以支持可复现性和未来研究。多视角评估显示,该代理在挑战样本上实现14.51%的WER降低,语义指标显著提升——MENLI提高7.59个百分点,槽位微观F1提升7.66个百分点。分析进一步表明,WER对领域迁移高度敏感,而语义指标与下游任务性能关联更紧密。
原文摘要 · Abstract (English)
While Automatic Speech Recognition (ASR) is typically benchmarked by word error rate (WER), real-world applications ultimately hinge on semantic fidelity. This mismatch is particularly problematic for dysarthric speech, where articulatory imprecision and disfluencies can cause severe semantic distortions. To bridge this gap, we introduce a Large Language Model (LLM)-based agent for post-ASR correction: a Judge-Editor over the top-k ASR hypotheses that keeps high-confidence spans, rewrites uncertain segments, and operates in both zero-shot and fine-tuned modes. In parallel, we release SAP-Hypo5, the largest benchmark for dysarthric speech correction, to enable reproducibility and future exploration. Under multi-perspective evaluation, our agent achieves a 14.51% WER reduction alongside substantial semantic gains, including a +7.59 pp improvement in MENLI and +7.66 pp in Slot Micro F1 on challenging samples. Our analysis further reveals that WER is highly sensitive to domain shift, whereas semantic metrics correlate more closely with downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。