用变异测试发现音频审核系统漏洞,能骗过主流平台检测。
Metamorphic Testing for Audio Content Moderation Software
- 设计变异测试框架MTAM,通过音调、噪声等微调生成避检音频
- 在5家商业平台和学术模型上测试,最高误检率达51.1%
- 可帮助安全团队发现审核系统薄弱点,适合内容安全研究者
音频平台如WhatsApp和Twitter的兴起改变了人们交流方式,但恶意内容(如仇恨言论、虚假广告、露骨内容)也日益泛滥,对心理健康造成严重影响。尽管已有多种音频内容审核工具,但攻击者可通过细微调整(如变调、加噪)规避检测,且现有工具对抗此类攻击的能力尚未充分评估。为此,本文提出MTAM——一种针对音频内容审核软件的变异测试框架。基于2000段音频样本,定义了14条变异关系,涵盖基于音频特征与启发式两类扰动。该框架生成仍具危害性但更易逃逸检测的测试用例。实验中,使用MTAM测试五家商业平台(Gladia、Assembly AI、Baidu、Nextdata、Tencent)及一个学术模型,结果表明其误检率(EFR)最高达38.6%、18.3%、35.1%、16.7%、51.1%,学术模型最高达45.7%。
原文摘要 · Abstract (English)
The rapid growth of audio-centric platforms and applications such as WhatsApp and Twitter has transformed the way people communicate and share audio content in modern society. However, these platforms are increasingly misused to disseminate harmful audio content, such as hate speech, deceptive advertisements, and explicit material, which can have significant negative consequences (e.g., detrimental effects on mental health). In response, researchers and practitioners have been actively developing and deploying audio content moderation tools to tackle this issue. Despite these efforts, malicious actors can bypass moderation systems by making subtle alterations to audio content, such as modifying pitch or inserting noise. Moreover, the effectiveness of modern audio moderation tools against such adversarial inputs remains insufficiently studied. To address these challenges, we propose MTAM, a Metamorphic Testing framework for Audio content Moderation software. Specifically, we conduct a pilot study on 2000 audio clips and define 14 metamorphic relations across two perturbation categories: Audio Features-Based and Heuristic perturbations. MTAM applies these metamorphic relations to toxic audio content to generate test cases that remain harmful while being more likely to evade detection. In our evaluation, we employ MTAM to test five commercial textual content moderation software and an academic model against three kinds of toxic content. The results show that MTAM achieves up to 38.6%, 18.3%, 35.1%, 16.7%, and 51.1% error finding rates (EFR) when testing commercial moderation software provided by Gladia, Assembly AI, Baidu, Nextdata, and Tencent, respectively, and it obtains up to 45.7% EFR when testing the state-of-the-art algorithms from the academy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。