构建首个面向真实场景的双语语音验证证据库,支持自然语言查询与可解释性分析。
SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification

- 用声学探针+大模型生成结构化语音档案,区分稳定特征与瞬时状态
- 包含10.2万说话人、56.7万条卡片、178万句级描述,支持跨模态检索与属性验证
- 适合语音验证、多模态对齐、可解释性研究者使用
现代说话人验证系统依赖有效但难以解释的说话人嵌入。现有语音-文本语料多用于可控合成或句子级描述,缺乏真实场景下说话人级别的监督信息。本文提出SpeakerCard-1M,一个基于VoxCeleb1/2与CN-Celeb1/2构建的双语证据基础说话人验证资源,其中"-1M"指释放中包含的178万句级描述。采用工具先行、大模型后置的方法:十种声学探针生成场域级证据,按结构化模式聚合为说话人档案(区分相对稳定特征与句级状态),再由受约束的大模型生成双语说话人卡片。数据集包含56.7万条说话人卡片记录、10.2万说话人、178万句级描述及说话人ID互斥的硬负样本三元组。我们定义两种面向说话人验证的跨模态协议:双向说话人-文本检索(T2S-R/S2T-R)和属性条件验证(AC-Verify)。在零样本强制选择设置下,对比双编码器基线与近期音频语言模型,联合音视频训练仅使VoxCeleb1-O的绝对错误等误率增加0.31%。在风格对称的对抗生成协议下,八款近期音频语言模型(参数量7B-30B+,含开源与闭源)在2分类强制选择任务中于音高层级的AC-Verify上得分49%-77%,远低于双编码器的88.66%。
原文摘要 · Abstract (English)
Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora target controllable synthesis or utterance-level captioning, offering limited speaker-level supervision for in-the-wild speaker recognition. This paper introduces SpeakerCard-1M, a bilingual speaker resource for evidence-grounded SV, derived from VoxCeleb1/2 and CN-Celeb1/2, where the ``-1M'' suffix refers to the 1.78M utterance-level captions contained in the release. We adopt a tool-first, LLM-last approach in which ten acoustic probes produce field-level evidence, the evidence is aggregated into speaker profiles under a schema that separates relatively stable traits from utterance-level states, and bilingual Speaker Cards are rendered by a constrained LLM that sees only the structured fields. The release includes 56.7k Speaker Card records over 10.2k speakers, 1.78M utterance-level captions, and speaker-ID-disjoint hard-negative triplets. We further define two SV-oriented cross-modal protocols, bidirectional Speaker-Text Retrieval (T2S-R / S2T-R) and Attribute-Conditioned Verification (AC-Verify), and compare a dual-encoder baseline against recent audio language models under a zero-shot forced-choice setting. Joint audio-text training costs only 0.31% absolute EER on VoxCeleb1-O relative to the audio-only baseline. Under a style-symmetric LLM-generated counterfactual protocol, eight recent audio language models (7B-30B+ parameters, both open- and closed-source) score 49-77% on pitch-level AC-Verify in a 2-way forced-choice setting, compared with 88.66% for our dual encoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。