让语音大模型能像人一样理解说话人身份与环境因素,支持可信的语音验证。
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

- 用分层语音标记器捕捉说话人多粒度特征,从整体音色到细微声学细节。
- 在单一语句中同时完成说话人画像、录音条件分析和对比判断,准确率提升显著。
- 生成可解释的决策链,适合需要可信语音验证的机器人和智能设备场景。
随着语音优先的AI代理在物理AI、对话机器人和无屏可穿戴设备中日益普及,语音大语言模型(audio-LLMs)必须整合说话人特异性理解,以支持用户授权、个性化及上下文感知交互。这需要建模是谁在说话、声音特征如何以及录音条件如何影响说话人线索。传统说话人验证系统提供强评分但缺乏语言证据,而现有音频-大模型和说话人感知语言模型仅能处理二元标签或描述性资料。我们提出SpeakerLLM,一种说话人专用的音频-大模型框架,统一单句话语说话人画像、录音条件理解、话语对说话人比较及证据组织化的验证推理,全部通过自然语言接口实现。我们构建了验证推理目标和决策组合策略,将画像级证据与最终‘相同/不同’判断分离,并将录音条件、画像证据与决策组织成结构化推理链。核心是分层说话人标记器,可捕获多粒度说话人证据:话语级嵌入总结身份与画像线索,帧级特征保留精细声学描述。实验表明,SpeakerLLM-Base在说话人画像与录音条件理解上优于通用音频-大模型;SpeakerLLM-VR保持高生成判决准确率,并生成基于监督验证推理模式的可解释决策链。我们将发布元数据增强的监督数据集和目标构建代码,以支持复现。
原文摘要 · Abstract (English)
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventional speaker verification systems provide strong scalar scores but little linguistic evidence, while current audio-LLMs and speaker-aware language models have limited ability to organize speaker information beyond binary labels or descriptive profiles. We present SpeakerLLM, a speaker-specialized audio-LLM framework that unifies single-utterance speaker profiling, recording-condition understanding, utterance-pair speaker comparison, and evidence-organized verification reasoning within a natural-language interface. We construct verification-reasoning targets and a decision-composition policy that separate profile-level evidence from the final same-or-different decision and organize recording condition, profile evidence, and the decision into a structured trace. At its core, SpeakerLLM uses a hierarchical speaker tokenizer designed to capture multiple granularities of speaker evidence. Utterance-level speaker embeddings summarize identity and profile-level cues, whereas frame-level speaker features preserve fine-grained acoustic descriptors. Experiments show that SpeakerLLM-Base improves speaker-profile and recording-condition understanding over general audio-LLMs, while SpeakerLLM-VR preserves strong generated-verdict accuracy and produces decision traces grounded in the supervised verification reasoning schema. We will release the metadata-enriched supervision dataset and target-construction code for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。