arXiv:2604.16749cs.SDcs.CL2026-04ACL被引 1

用对比引导的上下文学习提升音频伪造检测泛化能力

ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection

论文配图:ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
图 1 · 摘自论文原文
  • 通过成对比较推理引导语言模型识别并过滤无关声学特征
  • 在真实场景数据集上相较专用检测器宏F1提升最高2倍
  • 可适配开源语音大模型,支持无训练部署与结果解释

音频深度伪造构成重大安全威胁,但现有最先进检测系统难以泛化至真实环境中的伪造音频。本文提出一种新颖的上下文学习框架ICLAD,利用音频语言模型(ALM)实现对未见伪造音频的零训练泛化,并提供检测结果的文本依据。其核心是成对比较推理策略,引导ALM发现并剔除幻觉内容及与伪造无关的声学属性。ALM与专用检测器协同工作,路由机制将分布外样本送入ALM处理。在真实场景数据集上,ICLAD相较专用检测器的宏F1提升最高达2倍。进一步分析表明该方法具有灵活性,适用于近期开源音频语言模型的部署。

原文摘要 · Abstract (English)

Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel \textbf{I}n-\textbf{C}ontext \textbf{L}earning paradigm with comparison-guidance for \textbf{A}udio \textbf{D}eepfake detection (\textbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to $2\times$ relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.

音频伪造上下文学习大模型检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。