首个面向印地语的音视频深度伪造数据集,助力反虚假信息防御。
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset
- 用换脸、口型同步和语音克隆生成印地语多模态假视频。
- 现有检测模型在该数据集上准确率显著下降,挑战更大。
- 填补印地语深度伪造数据空白,适合本地化检测研究者使用。
深度伪造技术虽具创新潜力,但对隐私、信任与安全构成严重威胁。印度拥有庞大的印地语使用者群体,极易遭受深度伪造误导信息攻击。低数字素养的农村及半城市人群更易轻信视频内容。开发有效的防范框架需高质量、多样且大规模的数据集。现有主流数据集(如 FF-DF、DFDC)均基于英语,缺乏印地语支持。本文首次构建印地语音视频深度伪造数据集 HAV-DF,采用 faceswap、lipsyn 与 voice cloning 方法生成。该多步骤流程捕捉了印地语语音与面部表情的细微特征,为训练与评估印地语场景下的深度伪造检测模型提供坚实基础。此数据集为首个同时包含深度伪造视频与合成音频的多模态资源。相比其他知名数据集,其在 Headpose、Xception-c40 等检测方法下表现出更低的识别准确率,表明其更具挑战性,可能源于语言特性与多样化操纵手段。HAV-DF 填补了印地语深度伪造数据空白,有助于推动多语言深度伪造检测发展。
原文摘要 · Abstract (English)
Deepfakes offer great potential for innovation and creativity, but they also pose significant risks to privacy, trust, and security. With a vast Hindi-speaking population, India is particularly vulnerable to deepfake-driven misinformation campaigns. Fake videos or speeches in Hindi can have an enormous impact on rural and semi-urban communities, where digital literacy tends to be lower and people are more inclined to trust video content. The development of effective frameworks and detection tools to combat deepfake misuse requires high-quality, diverse, and extensive datasets. The existing popular datasets like FF-DF (FaceForensics++), and DFDC (DeepFake Detection Challenge) are based on English language.. Hence, this paper aims to create a first novel Hindi deep fake dataset, named ``Hindi audio-video-Deepfake'' (HAV-DF). The dataset has been generated using the faceswap, lipsyn and voice cloning methods. This multi-step process allows us to create a rich, varied dataset that captures the nuances of Hindi speech and facial expressions, providing a robust foundation for training and evaluating deepfake detection models in a Hindi language context. It is unique of its kind as all of the previous datasets contain either deepfake videos or synthesized audio. This type of deepfake dataset can be used for training a detector for both deepfake video and audio datasets. Notably, the newly introduced HAV-DF dataset demonstrates lower detection accuracy's across existing detection methods like Headpose, Xception-c40, etc. Compared to other well-known datasets FF-DF, and DFDC. This trend suggests that the HAV-DF dataset presents deeper challenges to detect, possibly due to its focus on Hindi language content and diverse manipulation techniques. The HAV-DF dataset fills the gap in Hindi-specific deepfake datasets, aiding multilingual deepfake detection development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。