arXiv:2505.15773eess.AScs.CL2025-05中稿 · INTERSPEECH 2025被引 6

构建首个中文语音毒性标注数据集,揭示语气中的隐性攻击

ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality

  • 构建13类话题的中文语音数据集,标注毒性类型与情绪来源
  • 融合声学语言情感特征,识别率超越纯文本模型
  • 适合语音安全、社交媒体监管方向研究者使用

尽管文本毒性检测研究广泛,但针对汉语口语的毒性分析仍存在空白。现有数据集缺乏对汉语特有的语调特征和文化表达的标注,导致语音毒性问题未被充分探索。为此,我们推出 ToxicTone——目前最大规模的公开中文语音毒性数据集,包含13个主题类别,涵盖辱骂、欺凌等毒性形式,以及愤怒、讽刺、轻蔑等情绪来源的详细标注。数据源自真实场景音频,反映自然交流状态。我们还提出一种多模态检测框架,结合先进语音与情绪编码器,融合声学、语言和情感特征。大量实验表明,该方法在识别隐藏毒性表达上显著优于仅基于文本的模型,证明语音特有线索对发现隐蔽攻击至关重要。

原文摘要 · Abstract (English)

Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique prosodic cues and culturally specific expressions in Mandarin leaves spoken toxicity underexplored. To address this, we introduce ToxicTone -- the largest public dataset of its kind -- featuring detailed annotations that distinguish both forms of toxicity (e.g., profanity, bullying) and sources of toxicity (e.g., anger, sarcasm, dismissiveness). Our data, sourced from diverse real-world audio and organized into 13 topical categories, mirrors authentic communication scenarios. We also propose a multimodal detection framework that integrates acoustic, linguistic, and emotional features using state-of-the-art speech and emotion encoders. Extensive experiments show our approach outperforms text-only and baseline models, underscoring the essential role of speech-specific cues in revealing hidden toxic expressions.

语音检测毒性识别中文数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。