用声音特征检测南印度语仇恨言论,验证多模态可行性
cantnlp@DravidianLangTech2025: A Bag-of-Sounds Approach to Multimodal Hate Speech Detection
- 将音频转为梅尔频谱图,采用'声音词袋'方法训练模型
- 在马拉雅拉姆语和泰米尔语上训练与开发集表现良好
- 证明文本与语音结合可实现多模态仇恨言论检测
本文报告了在第五届南印度语言语音、视觉与语言技术研讨会(DravidianLangTech-2025)上举行的南印度语言社交媒体数据分析共享任务(MSMDA-DL)的系统与结果。我们采用'声音词袋'方法,基于转换后的梅尔频谱图对语音(音频)数据进行仇恨言论检测模型训练。尽管候选模型在测试集上表现不佳,但在马拉雅拉姆语和泰米尔语的训练与开发阶段均展现出良好潜力。实验表明,在拥有充足且平衡的训练数据条件下,结合文本与语音数据构建多模态仇恨言论检测系统是可行的。
原文摘要 · Abstract (English)
This paper presents the systems and results for the Multimodal Social Media Data Analysis in Dravidian Languages (MSMDA-DL) shared task at the Fifth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages (DravidianLangTech-2025). We took a `bag-of-sounds' approach by training our hate speech detection system on the speech (audio) data using transformed Mel spectrogram measures. While our candidate model performed poorly on the test set, our approach offered promising results during training and development for Malayalam and Tamil. With sufficient and well-balanced training data, our results show that it is feasible to use both text and speech (audio) data in the development of multimodal hate speech detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。