arXiv:2412.11978cs.CLcs.SD2024-12中稿 · COLING 2025 main c…被引 1

用语音大模型自动验数据,省钱超40%还保质。

Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection

  • 用语音基础模型自动校验众包语音数据
  • 法德韩语数据实验中成本降40%以上
  • 适合需要大规模高质量语音数据的团队

尽管众包是语音数据采集的成熟方案,但非专业人士参与需额外质量控制。本文首次研究用语音基础模型(SFM)自动化验证流程,探索数据获取中的成本与质量权衡。在法语、德语和韩语数据上的实验表明,基于SFM的验证可显著减少对人工校验的依赖,预计成本降低超过40.0%且不损害最终数据质量。该成果为更高效、低成本、可扩展的语音数据采集提供了新路径。

原文摘要 · Abstract (English)

While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final data quality. To reduce the costs of these essential controls, this paper investigates the use of Speech Foundation Models (SFMs) to automate the validation process, examining for the first time the cost/quality trade-off in data acquisition. Experiments conducted on French, German, and Korean data demonstrate that SFM-based validation has the potential to reduce reliance on human validation, resulting in an estimated cost saving of over 40.0% without degrading final data quality. These findings open new opportunities for more efficient, cost-effective, and scalable speech data acquisition.

语音生成众包大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。