用少量数据实现10种印度语的音频暴力内容检测
Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning
- 基于预训练音频模型与元学习框架,实现跨语言低资源检测
- 50-200样本下仍保持有效识别,验证小样本可行性
- 适合多语言语音安全系统构建者参考
在线暴力内容检测在低资源语言和音频模态中仍研究不足。本文针对印度语场景,利用预训练音频表示(如Wav2Vec、Whisper)结合少样本学习(FSL)方法,在ADIMA数据集上探索跨语言暴力内容检测。通过将这些表示嵌入模型无关元学习(MAML)框架,实现对10种语言的滥用语言分类。实验采用50至200样本的多种样本量,评估数据稀缺对性能的影响。同时开展特征可视化分析以理解模型行为。研究表明,预训练模型在低资源情境下具备良好泛化能力,为多语言音频滥用检测提供了重要实践参考。
原文摘要 · Abstract (English)
Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource languages, in this case, in Indian languages using Few Shot Learning (FSL). Leveraging powerful representations from models such as Wav2Vec and Whisper, we explore cross-lingual abuse detection using the ADIMA dataset with FSL. Our approach integrates these representations within the Model-Agnostic Meta-Learning (MAML) framework to classify abusive language in 10 languages. We experiment with various shot sizes (50-200) evaluating the impact of limited data on performance. Additionally, a feature visualization study was conducted to better understand model behaviour. This study highlights the generalization ability of pre-trained models in low-resource scenarios and offers valuable insights into detecting abusive language in multilingual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。