AudioGuard为语音系统构建了跨威胁模型的全面安全防护机制。
AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
- 提出SoundGuard与ContentGuard双模块,分别检测音频原生危害与语义违规。
- 在多语言、儿童声音等12类风险场景中,准确率显著优于现有基线。
- 适用于语音助手、AI客服等需实时防护的语音交互系统。
音频正日益成为基础模型的主要交互接口,驱动实时语音助手应用。确保音频系统的安全性远超“说出不安全文本”范畴:真实风险可能源于音频原生有害声事件、说话人属性(如儿童声音)、语音伪造/克隆滥用,以及语音-内容组合性危害(如儿童声音+色情内容)。音频特性使得构建全面的评估基准或防护框架极具挑战。为此,我们对音频系统进行了大规模红队测试,系统性揭示其漏洞,构建了首个基于政策的综合性音频风险分类体系与AudioSafetyBench——首个覆盖多样化威胁模型的政策导向型音频安全基准。该基准支持多语言、可疑语音(如名人模仿、儿童声音)、高危语音-内容组合及非语音声事件。为应对这些威胁,我们提出AudioGuard统一防护框架,包含1)SoundGuard:波形级音频原生检测;2)ContentGuard:基于政策的语义保护。在AudioSafetyBench及四个互补基准上的大量实验表明,AudioGuard在保持显著更低延迟的同时,持续提升防护准确率,优于强基线音频大模型方案。
原文摘要 · Abstract (English)
Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants. Ensuring safety in audio systems is inherently more complex than just "unsafe text spoken aloud": real-world risks can hinge on audio-native harmful sound events, speaker attributes (e.g., child voice), impersonation/voice-cloning misuse, and voice-content compositional harms, such as child voice plus sexual content. The nature of audio makes it challenging to develop comprehensive benchmarks or guardrails against this unique risk landscape. To close this gap, we conduct large-scale red teaming on audio systems, systematically uncover vulnerabilities in audio, and develop a comprehensive, policy-grounded audio risk taxonomy and AudioSafetyBench, the first policy-based audio safety benchmark across diverse threat models. AudioSafetyBench supports diverse languages, suspicious voices (e.g., celebrity/impersonation and child voice), risky voice-content combinations, and non-speech sound events. To defend against these threats, we propose AudioGuard, a unified guardrail consisting of 1) SoundGuard for waveform-level audio-native detection and 2) ContentGuard for policy-grounded semantic protection. Extensive experiments on AudioSafetyBench and four complementary benchmarks show that AudioGuard consistently improves guardrail accuracy over strong audio-LLM-based baselines with substantially lower latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。