构建隐私保护的异常检测数据集,用文本描述视频中的侵权行为
PV-VTT: A Privacy-Centric Dataset for Mission-Specific Anomaly Detection and Natural Language Interpretation
- 用图神经网络生成视频描述提示,结合单帧图像与文本
- 仅提供视频特征向量,不泄露原始视频数据
- 适合关注隐私保护与可解释性分析的研究者
视频犯罪检测是计算机视觉与人工智能的重要应用。现有数据集多聚焦于完整视频中严重犯罪的识别,常忽略可能预防犯罪的前兆行为(即隐私侵犯)。为此,我们提出PV-VTT(Privacy Violation Video To Text)——一个专注于识别隐私侵犯的多模态数据集,提供视频与文本的详细标注。为保障视频中个体隐私,仅提供视频特征向量,不公开原始视频数据。鉴于隐私侵犯具有模糊性和上下文依赖性,我们设计基于图神经网络(GNN)的视频描述模型,通过单帧图像与相关文本生成适配大语言模型(LLM)的提示,降低输入令牌数,实现高性价比、高质量的视频描述生成。大量实验验证了该方法在视频描述任务中的有效性与可解释性,以及PV-VTT数据集的灵活性。
原文摘要 · Abstract (English)
Video crime detection is a significant application of computer vision and artificial intelligence. However, existing datasets primarily focus on detecting severe crimes by analyzing entire video clips, often neglecting the precursor activities (i.e., privacy violations) that could potentially prevent these crimes. To address this limitation, we present PV-VTT (Privacy Violation Video To Text), a unique multimodal dataset aimed at identifying privacy violations. PV-VTT provides detailed annotations for both video and text in scenarios. To ensure the privacy of individuals in the videos, we only provide video feature vectors, avoiding the release of any raw video data. This privacy-focused approach allows researchers to use the dataset while protecting participant confidentiality. Recognizing that privacy violations are often ambiguous and context-dependent, we propose a Graph Neural Network (GNN)-based video description model. Our model generates a GNN-based prompt with image for Large Language Model (LLM), which deliver cost-effective and high-quality video descriptions. By leveraging a single video frame along with relevant text, our method reduces the number of input tokens required, maintaining descriptive quality while optimizing LLM API-usage. Extensive experiments validate the effectiveness and interpretability of our approach in video description tasks and flexibility of our PV-VTT dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。