区分视频真假已不够,新基准可识别伪造视频背后的23种意图。
Beyond Real versus Fake Towards Intent-Aware Video Analysis
- 构建包含5168个视频的意图分析基准,标注23类真实动机。
- 多模态模型融合时空特征、音频与文本,准确识别伪造视频目的。
- 适合安全、传媒与政策研究者,关注内容深层意图而非真假。
生成模型的快速发展使得深度伪造视频日益逼真,带来重大社会与安全风险。现有检测方法仅关注真实与伪造的区分,却无法回答核心问题:伪造视频背后的意图是什么?为此,我们提出IntentHQ——一个以人类为中心的意图分析新基准,推动从真实性验证转向对视频上下文的深入理解。IntentHQ包含5168个精心采集并标注的视频,涵盖23种细粒度意图类别,如“金融诈骗”、“间接营销”、“政治宣传”和“制造恐慌”。我们采用监督与自监督的多模态模型,整合时序视频特征、音频处理与文本分析,推断视频背后的真实动机与目标。所提模型结构精简,能有效区分多种意图类别。
原文摘要 · Abstract (English)
The rapid advancement of generative models has led to increasingly realistic deepfake videos, posing significant societal and security risks. While existing detection methods focus on distinguishing real from fake videos, such approaches fail to address a fundamental question: What is the intent behind a manipulated video? Towards addressing this question, we introduce IntentHQ: a new benchmark for human-centered intent analysis, shifting the paradigm from authenticity verification to contextual understanding of videos. IntentHQ consists of 5168 videos that have been meticulously collected and annotated with 23 fine-grained intent-categories, including "Financial fraud", "Indirect marketing", "Political propaganda", as well as "Fear mongering". We perform intent recognition with supervised and self-supervised multi-modality models that integrate spatio-temporal video features, audio processing, and text analysis to infer underlying motivations and goals behind videos. Our proposed model is streamlined to differentiate between a wide range of intent-categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。