arXiv:2605.02223cs.SDcs.CV2026-05

针对语音篡改中多区域局部替换难题,提出新数据集、检测框架与评估指标。

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

论文配图:Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
图 1 · 摘自论文原文
  • 采用大语言模型引导语义替换与神经语音克隆生成多区域篡改语音
  • 在仅2-7%内容被篡改时仍能精准定位多个未知数量的伪造片段
  • 首次实现无需预知篡改段数的细粒度语音伪造定位,适合安全与司法场景

语音克隆与文本转语音技术的进步使得部分语音篡改——即在一段话中替换少量词语以改变语义但保持说话人特征——成为现实威胁。现有音频深度伪造检测基准多聚焦于整句二分类或单区域篡改,难以应对多个未知数量的篡改区域。为此,本文提出三项贡献:第一,构建MIST(Multiregion Inpainting Speech Tampering)数据集,涵盖6种语言,每条语音含1-3个独立的词级篡改区域,通过大语言模型引导语义替换和神经语音克隆生成,伪造内容占比仅2-7%;第二,提出ISA(Iterative Segment Analysis)框架,一种不依赖主干网络的粗到精滑动窗口分类方法,具备间隙容忍的区域提案与边界精修能力,可无先验知识恢复所有篡改区域;第三,定义SF1@tau,基于时间交并比匹配的段级F1指标,联合评估区域数量准确性和定位精度。零样本测试表明,现有深度伪造检测器对词粒度局部篡改几乎无效:训练于全合成语音的整句分类器对仅2-7%被篡改的MIST语音赋予接近零的伪造概率。ISA在该挑战性场景中持续优于非迭代基线,相关数据集、代码与评估工具包已公开发布。

原文摘要 · Abstract (English)

Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identity - an increasingly realistic threat. Existing audio deepfake detection benchmarks focus on utterance-level binary classification or single-region tampering, leaving a critical gap in detecting and localizing multiple inpainted segments whose count is unknown a priori. We address this gap with three contributions. First, we introduce MIST (Multiregion Inpainting Speech Tampering), a large-scale multilingual dataset spanning 6 languages with 1-3 independently inpainted word-level segments per utterance, generated via LLM-guided semantic replacement and neural voice cloning, with fake content constituting only 2-7% of each utterance. Second, we propose ISA (Iterative Segment Analysis), a backbone-agnostic framework that performs coarse-to-fine sliding-window classification with gap-tolerant region proposal and boundary refinement to recover all tampered regions without prior knowledge of their count. Third, we define SF1@tau, a segment-level F1 metric based on temporal IoU matching that jointly evaluates region count accuracy and localization precision. Zero-shot evaluation reveals that partial inpainting at word granularity remains unsolved by existing deepfake detectors: utterance-level classifiers trained on fully synthesized speech assign near zero fake probability to MIST utterances where only 2-7% of content is manipulated. ISA consistently outperforms non-iterative baselines in this challenging setting, and the dataset, code, and evaluation toolkit are publicly released.

语音伪造细粒度定位检测基准多区域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。