通过音视频对齐提升儿童有害内容细粒度检测效果
SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
- 用级联跨模态变换器实现音视频特征精准对齐
- 在HatefulClips数据集上达到91.3%的检测准确率,优于单模态和基础融合方法
- 适合需要高精度内容审核的平台或研究者参考
随着视频分享平台的发展,儿童观众数量激增,对暴力、露骨等有害内容的精准检测需求日益迫切。恶意用户常通过在极少数帧中嵌入不安全内容来规避检测。尽管已有研究聚焦视觉线索并提升了细粒度检测能力,音频特征仍被严重忽视。本文提出SNIFR框架,将音频与视觉信息融合用于细粒度儿童有害内容检测。该框架采用变换器编码器进行模态内交互,再通过级联跨模态变换器实现模态间对齐。实验表明,该方法在性能上显著超越单模态及基线融合方法,达到新的最优水平。
原文摘要 · Abstract (English)
As video-sharing platforms have grown over the past decade, child viewership has surged, increasing the need for precise detection of harmful content like violence or explicit scenes. Malicious users exploit moderation systems by embedding unsafe content in minimal frames to evade detection. While prior research has focused on visual cues and advanced such fine-grained detection, audio features remain underexplored. In this study, we embed audio cues with visual for fine-grained child harmful content detection and introduce SNIFR, a novel framework for effective alignment. SNIFR employs a transformer encoder for intra-modality interaction, followed by a cascaded cross-transformer for inter-modality alignment. Our approach achieves superior performance over unimodal and baseline fusion methods, setting a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。