构建首个面向手机录音的模型识别数据集,助力音视频溯源研究。
POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
- 收集20款新型手机在受控环境下的多类音频数据。
- 基于最新分类器,在该数据集上实现高精度手机型号识别。
- 适合数字取证、媒体可信度验证等领域的研究人员使用。
多媒体内容的来源归属是法证分析中的关键挑战,旨在确定内容的生成设备,对法律程序和完整性调查具有重要价值。已有研究涵盖从相机型号识别到语音合成器或麦克风型号检测等多个领域。近年来,机器学习驱动的方法显著优于传统信号处理技术,但其依赖大量训练数据,且数据需紧跟技术发展。然而,智能手机快速迭代导致现有数据集往往过时或采集不一致,限制了模型在真实场景中的有效性。本文提出POLIPHONE数据集,包含20款最新智能手机在受控环境下录制的音频,涵盖语音、音乐和环境声等多种类型,确保可复现性和可扩展性。我们还使用先进分类器对数据集进行基准测试,验证其在手机型号识别任务中的实用性。
原文摘要 · Abstract (English)
When dealing with multimedia data, source attribution is a key challenge from a forensic perspective. This task aims to determine how a given content was captured, providing valuable insights for various applications, including legal proceedings and integrity investigations. The source attribution problem has been addressed in different domains, from identifying the camera model used to capture specific photographs to detecting the synthetic speech generator or microphone model used to create or record given audio tracks. Recent advancements in this area rely heavily on machine learning and data-driven techniques, which often outperform traditional signal processing-based methods. However, a drawback of these systems is their need for large volumes of training data, which must reflect the latest technological trends to produce accurate and reliable predictions. This presents a significant challenge, as the rapid pace of technological progress makes it difficult to maintain datasets that are up-to-date with real-world conditions. For instance, in the task of smartphone model identification from audio recordings, the available datasets are often outdated or acquired inconsistently, making it difficult to develop solutions that are valid beyond a research environment. In this paper we present POLIPHONE, a dataset for smartphone model identification from audio recordings. It includes data from 20 recent smartphones recorded in a controlled environment to ensure reproducibility and scalability for future research. The released tracks contain audio data from various domains (i.e., speech, music, environmental sounds), making the corpus versatile and applicable to a wide range of use cases. We also present numerous experiments to benchmark the proposed dataset using a state-of-the-art classifier for smartphone model identification from audio recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。