arXiv:2505.00409eess.AScs.AI2025-05

自动匿名化会显著影响病理语音的感知质量,但临床评估仍可保持准确。

Perceptual implications of automatic anonymization in pathological speech

  • 通过人类听觉实验评估自动匿名化对病理语音的影响
  • 91%听众能识别匿名语音,质量评分下降30分(0-100)
  • 临床严重度评级几乎不受影响,适合临床使用验证

自动匿名化日益用于伦理共享临床语音数据,但其感知与临床影响尚不明确。本研究采用结构化评估协议,邀请10名德语母语及非母语者(涵盖临床与信号处理背景)对180名来自口颌裂、运动性构音障碍、发音障碍、发声障碍及成人/儿童对照组的德语说话者原始录音及其自动匿名化版本进行四项任务评估:零样本图灵式辨别、简短熟悉后的少样本辨别、五级质量评分、以及资深耳鼻喉科医生的四级盲评临床严重度。听者在零样本下以91%准确率、少样本下93%准确率检测到匿名化,不同病症间差异显著(p=0.008),且熟悉后该差异减弱。感知质量在0-100量表上平均下降30个百分点(p<0.001),并重塑了各组间的感知质量排序。母语影响检测能力但不影响质量下降;领域专长影响质量评价但不影响检测能力,呈现双分离效应;性别与年龄无显著偏差。临床严重度评分在运动性构音障碍、发音障碍和发声障碍中保持近完美一致性(加权克朗巴赫系数κ=0.87–0.94),无录音跨等级变化。关键发现:感知结果与标准计算隐私指标脱钩——计算匿名化最强的病症反而感知最不明显,反之亦然。研究呼吁建立按疾病分层、听者分层、临床医生验证的最小评估标准,方可批准匿名语音用于临床应用。

原文摘要 · Abstract (English)

Automatic anonymization is increasingly used to enable ethical sharing of clinical speech, yet its perceptual and clinical consequences remain undercharacterized. We present a human-centered evaluation of automatically anonymized pathological speech, using a structured protocol with ten native and non-native German listeners spanning clinical and signal-processing expertise. The cohort comprised 180 German speakers from CLP, Dysarthria, Dysglossia, Dysphonia, and adult and child controls. Each original recording and its automatically-anonymized counterpart was evaluated on four tasks: zero-shot Turing-style discrimination, few-shot discrimination after brief familiarization, 5-point quality rating, and 4-point blinded clinical severity rating by a senior phoniatrician. Listeners detected anonymization at 91% zero-shot and 93% few-shot accuracy, with significant variation across disorders (p=0.008) that attenuated with familiarization. Perceived quality dropped by 30 ppts on a 0-100 scale (p<0.001), reorganizing the perceived-quality hierarchy across groups. Native language modulated detectability but not quality degradation, while domain expertise modulated quality degradation but not detectability, a double dissociation between the two listener attributes; speaker sex and age produced no detectable bias. Clinical severity ratings were preserved at near-perfect agreement in Dysarthria, Dysglossia, and Dysphonia (quadratic-weighted Cohen's kappa 0.87-0.94), with no recording shifting by more than one grade. Crucially, perceptual outcomes were decoupled from the standard computational privacy metric: the pathology with the strongest computational anonymization was the least perceptually conspicuous, and vice versa. These findings argue for disorder-stratified, listener-stratified, clinician-validated evaluation as the minimum standard for licensing anonymized speech for clinical use.

语音匿名化临床评估感知质量病理语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。