提出噪声压力测试,发现语音清晰度与医疗安全不匹配。
Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

- 同一对话注入不同噪声,固定模型参数,隔离噪声影响。
- 静音背景噪声使错误率仅升0.71个百分点,但不安全输出翻倍。
- 小噪声扰动可改变临床意义却不显著提高错误率,适合医疗AI评估。
环境临床记录系统越来越多地结合自动语音识别与大语言模型来实现文档自动化。然而,传统指标如词错误率会掩盖系统性安全退化问题。本文提出一种成对声学压力测试,以分离噪声对临床推理的因果影响。在相同对话中,我们注入多种噪声类型,同时保持下游模型配置不变。关键发现是:信号保真度与临床安全性之间存在危险脱节。静音背景噪声仅使词错误率提升0.71个百分点,却几乎使不安全输出率翻倍。分析表明,微小的声学扰动可在不显著增加错误率的情况下反转临床含义。此外,我们展示了一种轻量级缓解策略,在无需模型微调的前提下有效降低噪声下的安全退化。
原文摘要 · Abstract (English)
Ambient clinical scribes increasingly combine Automatic Speech Recognition with Large Language Models to automate documentation. However, traditional metrics like Word Error Rate mask systemic safety degradation. We present a paired acoustic stress test to isolate the causal impact of noise on clinical reasoning. For the same dialogues, we inject diverse noise types while keeping the downstream model configuration frozen. Crucially, we uncover a dangerous disconnect between signal fidelity and clinical safety. Stationary ambient noise increased the Word Error Rate by a negligible 0.71 percentage points yet nearly doubled the rate of unsafe outputs. Our analysis reveals that minor acoustic perturbations can invert clinical meaning without substantially inflating error rates. Furthermore, we demonstrate a lightweight mitigation strategy that mitigates safety degradation under noisy conditions without requiring model fine tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。