arXiv:2607.14753cs.SDcs.AI2026-07

用大模型提升语音验证防伪造能力,支持可解释审计。

Large Audio Language Models for Spoofing-Aware Speaker Verification

论文配图:Large Audio Language Models for Spoofing-Aware Speaker Verification
图 1 · 摘自论文原文
  • 用大音频语言模型实现统一的防伪语音验证
  • 适应训练后性能媲美传统流水线系统
  • 生成自然语言理由,适合需要透明性的场景

近期文本转语音和语音克隆技术的进步使高质量伪造语音变得廉价且可扩展,严重威胁语音认证系统,尤其是自动说话人验证(ASV)。现有防御方法主要采用二元反制措施(CM)或防伪说话人验证(SASV),当前系统以模块化融合和级联流水线为主。尽管大音频语言模型(LALMs)在相关音频任务(如CM和ASV)中表现优异,但其在SASV中的应用尚未探索,尽管其具备生成自然语言推理、超越判别性预测的鲁棒性。本文系统评估了LALMs在零样本提示、监督微调、基于推理的训练及强化学习优化下的SASV性能。结果表明,预训练的LALMs在零样本设置下仅接近随机水平,证实其并非天生适合SASV;但通过任务特定适应可显著缩小差距。进一步发现,多种不同路径均可达到竞争性性能。这些结果将LALMs定位为统一SASV的有前景且可审计基础,同时明确了传统级联系统仍具优势的领域。

原文摘要 · Abstract (English)

Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.

语音验证防伪检测大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。