arXiv:2510.02181cs.HCcs.AI2025-10被引 1

让听人实时纠错,帮聋哑人士自动生成语音识别模型

EvolveCaptions: Empowering DHH Users Through Real-Time Collaborative Captioning

  • 听人现场纠正错误,系统生成语音提示让聋哑人录制训练
  • 一小时内平均仅用5分钟录音,所有用户WER显著下降
  • 适合需要实时无障碍沟通的聋哑人群体使用

自动语音识别(ASR)系统在转录聋哑人士语音时经常出错,尤其在实时对话中。现有个性化方法通常需要大量预先录制的数据,并将适应负担留给聋哑说话者。我们提出 EvolveCaptions,一个支持现场个性化、低负担的实时协作式 ASR 适应系统。听人参与者在实时对话中纠正识别错误,系统据此生成短时、音素针对性的提示,供聋哑用户录制,用于微调 ASR 模型。在12名聋哑用户与6名听力正常参与者参与的实验中,系统在使用一小时内将所有聋哑用户的词错误率(WER)降低,平均仅需5分钟录音时间。参与者认为系统直观、低耗且易于融入交流。这些结果表明,协作式实时 ASR 适应具有提升沟通公平性的潜力。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems often fail to accurately transcribe speech from Deaf and Hard of Hearing (DHH) individuals, especially during real-time conversations. Existing personalization approaches typically require extensive pre-recorded data and place the burden of adaptation on the DHH speaker. We present EvolveCaptions, a real-time, collaborative ASR adaptation system that supports in-situ personalization with minimal effort. Hearing participants correct ASR errors during live conversations. Based on these corrections, the system generates short, phonetically targeted prompts for the DHH speaker to record, which are then used to fine-tune the ASR model. In a study with 12 DHH and six hearing participants, EvolveCaptions reduced Word Error Rate (WER) across all DHH users within one hour of use, using only five minutes of recording time on average. Participants described the system as intuitive, low-effort, and well-integrated into communication. These findings demonstrate the promise of collaborative, real-time ASR adaptation for more equitable communication.

语音识别聋哑辅助实时交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。