首个多模态音频视频生成安全基准,揭示组合风险与检测盲区。
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

- 构建11种跨模态组合的攻击实例库,覆盖全部非单模态配置
- 发现现有安全防护在组合输入下失效,有害语义可由无害输入合成
- 适合研究多模态生成安全、模型鲁棒性与防御机制的开发者
音频视频生成正从提示驱动转向多模态条件生成,文本、图像、音频与视频可共同决定输出结果。这一转变改变了安全评估的本质:有害意图可能不再源自单一输入,而是源于各模态间及时间维度上的交互作用。然而,现有安全评测仍以提示为中心或依赖固定输入接口,难以系统研究此类组合性风险。为此,我们提出 Multi2AV-Safety,据我们所知,首个覆盖所有11种非单模态 T/I/A/V 条件组合的音频视频生成安全基准,包含11,024个攻击实例。在该基准上的评估揭示了代表性多模态安全防护在多种攻击机制和危害证据结构下的系统性弱点。结果表明存在两种互补的失效模式:有害语义可由单独无害的输入组合产生;而明确的有害信号在融入良性多模态上下文后更难被检测。这些发现指明了保障多模态条件音频视频生成中的核心能力缺口——当前安全机制无法可靠地跨模态与时间整合安全证据,即使所有输入均可观测。数据集将于2026年10月公开发布。
原文摘要 · Abstract (English)
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify \emph{compositional risk perception} as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。