直接用音频检测印地语系仇恨言论,效果接近全监督模型。
Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages

- 用CLAP模型直接分析音频,跳过不靠谱的语音转写。
- 无需微调仅用少量标注数据,性能已逼近全监督系统。
- 适合低资源语言的快速部署,尤其适合数据稀缺场景。
仇恨言论正越来越多以语音形式出现,如语音消息、通话和短视频。现有系统多依赖语音转写后再分类,但在缺乏强语音识别器的语言中转写不可靠,且丢失了语气与情绪等关键信息。本文研究是否可直接通过音频检测仇恨言论,采用学习声语联合表征的CLAP模型,在ADIMA数据集涵盖的十种印地语系语言上进行评估。仅用现成的CLAP音频表示训练轻量级分类器,无需微调模型,性能即达到全监督系统的1-3个百分点以内,远超零样本提示。进一步用每语言少量标注样本微调,收益有限且表现不稳定。结果表明,基于CLAP的音频表示已为多语言仇恨言论检测提供低成本强基础,显著降低实际应用中的标注需求。
原文摘要 · Abstract (English)
Abusive and hateful speech is increasingly spoken rather than written, surfacing in voice notes, calls, and short-form videos. Most detection systems still transcribe speech to text before classifying it, but transcription is unreliable for languages lacking strong speech recognisers, and it discards the tone and emotion that often carry the abuse itself. This paper examines whether abusive speech can instead be detected directly from audio, using CLAP, a model that learns a shared representation of sound and language, evaluated across ten Indic languages in the ADIMA dataset. A lightweight classifier trained on CLAP's existing audio representations, without adapting the model itself, comes within one to three points of a fully supervised system, and far outperforms prompting with no labelled examples at all. Further adaptation with a handful of labelled examples per language yields little extra benefit, varying unpredictably across languages. CLAP-based audio representations thus already offer a strong, inexpensive foundation for detecting abusive speech across languages, lowering the labelled data needed in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。