arXiv:2509.21554cs.CLcs.AI2025-09被引 1

研究非洲口音英语的说话人分离,发现临床对话误差显著更高。

Domain-Aware Speaker Diarization On African-Accented English

  • 在重叠语音场景下评估多个系统,发现临床对话误差更大。
  • 临床对话错误主要来自误报和漏检,与短句和频繁重叠有关。
  • 轻量级适配可降错但难根治,适合资源有限团队复现。

本研究考察非洲口音英语在说话人分离中的领域影响。我们在通用和临床对话上,采用严格重叠评分的DER协议评估多个生成与开源系统。结果显示,临床语音始终存在显著领域惩罚,且跨模型一致。误差分析表明,该现象主要源于误报和漏检,与短时发言及频繁重叠相关。通过在口音匹配数据上微调分割模块进行轻量级领域适配,虽能降低误差但无法消除差距。贡献包括跨领域可控基准、简洁的误差分解与对话级分析方法,以及易于复现的适配方案。结果建议未来应关注重叠感知分割与临床资源均衡配置。

原文摘要 · Abstract (English)

This study examines domain effects in speaker diarization for African-accented English. We evaluate multiple production and open systems on general and clinical dialogues under a strict DER protocol that scores overlap. A consistent domain penalty appears for clinical speech and remains significant across models. Error analysis attributes much of this penalty to false alarms and missed detections, aligning with short turns and frequent overlap. We test lightweight domain adaptation by fine-tuning a segmentation module on accent-matched data; it reduces error but does not eliminate the gap. Our contributions include a controlled benchmark across domains, a concise approach to error decomposition and conversation-level profiling, and an adaptation recipe that is easy to reproduce. Results point to overlap-aware segmentation and balanced clinical resources as practical next steps.

说话人分离非洲口音临床语音误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。