用大模型检测语音中被篡改的词语,发现它依赖特定编辑模式。
Can LLMs Help Localize Fake Words in Partially Fake Speech?
- 通过下一词预测构建语音大模型定位篡改词
- 在两个数据集上依赖词级情感替换等编辑模式
- 适合研究伪造语音检测与模型泛化能力的学者
大型语言模型(LLMs)在大规模文本训练后,在多项任务中表现强劲。受此启发,本文探究文本训练的LLM能否帮助定位部分伪造语音中的虚假词语(仅部分词语被修改)。我们构建了一个语音LLM,通过下一词预测实现虚假词语定位。在AV-Deepfake1M和PartialEdit数据集上的实验与分析表明,该模型常依赖训练数据中学到的编辑风格模式,特别是针对这两个数据库中常见的词级极性替换作为定位线索。尽管此类特定模式在域内场景下有效,但如何避免对特定模式的过度依赖并提升对未见编辑风格的泛化能力,仍是开放问题。
原文摘要 · Abstract (English)
Large language models (LLMs), trained on large-scale text, have recently attracted significant attention for their strong performance across many tasks. Motivated by this, we investigate whether a text-trained LLM can help localize fake words in partially fake speech, where only specific words within a speech are edited. We build a speech LLM to perform fake word localization via next token prediction. Experiments and analyses on AV-Deepfake1M and PartialEdit indicates that the model frequently leverages editing-style pattern learned from the training data, particularly word-level polarity substitutions for those two databases we discussed, as cues for localizing fake words. Although such particular patterns provide useful information in an in-domain scenario, how to avoid over-reliance on such particular pattern and improve generalization to unseen editing styles remains an open question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。