arXiv:2509.19755cs.SDeess.AS2025-09被引 7

用大模型做语音身份验证,效果有潜力但需调优。

Can Audio Large Language Models Verify Speaker Identity?

  • 把语音验证转为音频问答任务,零样本测试表现有限。
  • 微调后性能显著提升,但仍不及传统模型。
  • 结合说话内容验证,效果接近串行系统,适合统一设计。

本文研究将音频大语言模型(ALLM)用于语音身份验证(SV)。我们将SV重新定义为音频问答任务,在公开基准上进行零样本评估,发现当前ALLM在零样本下的SV能力有限,且在多变声学条件下常表现不佳。为此,我们在语音验证数据上进行监督微调,并提出基于规则的困难样本采样策略以构建更具挑战性的训练对。轻量级微调显著提升了性能,尽管与传统模型仍存在差距。随后,我们扩展至文本依赖型语音验证,联合查询ALLM以验证说话人身份和语音内容,结果达到与串行语音识别-语音验证系统相当的水平。研究结果表明,经适当适配后,ALLM具备作为鲁棒语音验证系统的统一模型的潜力,同时保持通用音频理解能力。

原文摘要 · Abstract (English)

This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive zero-shot evaluations on public benchmarks, showing that current ALLMs have limited zero-shot SV capability and often struggle in diverse acoustic conditions. To address this challenge, we perform supervised fine-tuning on speaker verification data. A rule-based hard pair sampling strategy is proposed to construct more challenging training pairs. Lightweight fine-tuning substantially improves the performance, though there is still a gap between ALLMs and conventional models. Then, we extend to text-dependent SV by jointly querying ALLMs to verify speaker identity and spoken content, yielding results competitive with cascaded ASR-SV systems. Our findings demonstrate that with proper adaptation, ALLMs hold substantial potential as a unified model for robust speaker verification systems, while maintaining the general audio understanding capabilities.

语音验证大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。