arXiv:2506.00636cs.CL2025-06中稿 · presentation at IN…

首个越南语语音毒性片段检测数据集,支持精准识别网络语音中的不当内容。

ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances

  • 构建语音转文本+毒性检测流水线,实现细粒度毒性内容定位。
  • 在11,000个音频样本上,微调后ASR模型词错误率显著降低。
  • 为低资源语言语音安全研究提供首个可复现的基准数据集。

在线平台中的不当言论日益引发关注,影响用户体验与网络安全。尽管文本毒性检测已较为成熟,但针对低资源语言如越南语的语音毒性检测仍处于空白。本文提出首个越南语语音毒性片段检测数据集ViToSA,包含11,000个音频样本(总计25小时),并配有精确的人工标注转录文本。我们设计了一个结合自动语音识别(ASR)与毒性片段检测(TSD)的处理流程,实现对语音中毒性内容的细粒度识别。实验表明,在ViToSA上微调的ASR模型在转录毒性语音时词错误率(WER)显著下降;同时,基于文本的毒性片段检测模型优于现有基线方法。本工作建立了越南语语音毒性检测的新基准,为未来语音内容审核研究铺平道路。

原文摘要 · Abstract (English)

Toxic speech on online platforms is a growing concern, impacting user experience and online safety. While text-based toxicity detection is well-studied, audio-based approaches remain underexplored, especially for low-resource languages like Vietnamese. This paper introduces ViToSA (Vietnamese Toxic Spans Audio), the first dataset for toxic spans detection in Vietnamese speech, comprising 11,000 audio samples (25 hours) with accurate human-annotated transcripts. We propose a pipeline that combines ASR and toxic spans detection for fine-grained identification of toxic content. Our experiments show that fine-tuning ASR models on ViToSA significantly reduces WER when transcribing toxic speech, while the text-based toxic spans detection (TSD) models outperform existing baselines. These findings establish a novel benchmark for Vietnamese audio-based toxic spans detection, paving the way for future research in speech content moderation.

语音检测毒性内容越南语ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。