arXiv:2606.01704eess.AS2026-06

用语音验证亲缘关系,发现声纹能反映家族线索。

Kinship Verification Using Voice

论文配图:Kinship Verification Using Voice
图 1 · 摘自论文原文
  • 构建新评估协议,避免数据泄露,支持开放集验证
  • 零样本验证中最低误判率20.8%,严格条件下升至39.7%
  • 提出异构处理策略缓解年龄差异影响,适合家庭关系分析研究

基于语音的亲缘关系验证(Kinship Verification, KV)任务旨在判断两人是否具有生物学亲属关系,目前关注度较低。本文以大规模音视频数据集KAN-AV为基础,提出改进的评估协议,控制多种混杂因素,并采用家族互斥的训练-测试划分方式,应对开放集场景下的KV问题。我们分析了说话人验证与KV之间的密切关联,发现亲属相似性在两类任务中起相反作用。进一步地,使用三种神经声纹提取器(ECAPA-TDNN、WavLM-ECAPA、ReDimNet)结合不同后端方法进行实验。在包含同说话人目标试验的零样本设置下,ReDimNet达到最低等错误率(EER)20.8%;但在排除同说话人目标试验的严格亲缘验证中,性能下降至39.7%。最优可训练后端通过异构处理嵌入对,减轻年龄差异影响,获得EER 32.0%(含说话人目标试验时为18.6%)。结果表明:尽管挑战大,但声纹嵌入确实编码了家族线索,为语音驱动的亲缘分析奠定基础。

原文摘要 · Abstract (English)

Kinship verification (KV) from voice, the task of determining whether two speakers are biologically related, has received only little attention. Our work establishes a foundational basis for this emerging frontier, contributing to both performance evaluation and detection methodologies. First, leveraging the speech recordings of the large-scale audio-visual dataset, KAN-AV, we propose a revised evaluation protocol that controls for various confounders and adopts a family-disjoint train--test split to address open-set KV. Second, we analyze the close connection between speaker verification and KV, showing that genealogical similarity of speaker pairs plays opposite roles in the two tasks. Third, we tackle KV using three neural speaker embedding extractors (ECAPA-TDNN, WavLM-ECAPA, and ReDimNet) combined with various back-ends. In zero-shot KV including same-speaker target trials, ReDimNet achieves the lowest equal error rate (EER) of $20.8\%$; however, performance degrades to $39.7\%$ under strict kin trials, where same-speaker target trials are excluded. Our best trainable back-end, which applies asymmetric processing of the embedding pair to mitigate age-difference effects, obtains an EER of $32.0\%$ ($18.6\%$ with speaker target trials included). These results highlight the difficulty of KV while showing that speaker embeddings encode familial cues, offering a promising foundation for voice-based kinship analysis.

亲缘识别声纹分析语音验证嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。