用语音年龄转换提升亲子关系识别准确率
Audio-based Kinship Verification Using Age Domain Conversion
- 通过优化的CycleGAN-VC3实现语音年龄转换,构建标准化年龄域
- 在KAN_AV数据集上准确率显著提升,验证了方法有效性
- 适合音频识别、生物特征认证等领域的研究人员参考
基于语音的亲属关系验证(AKV)在家庭安防监控、法医鉴定和社会网络分析中具有重要意义。其主要挑战源于不同个体间样本年龄差异带来的域偏移问题。为此,本文提出‘年龄标准化域’概念,利用优化的CycleGAN-VC3网络进行语音年龄转换,生成同龄域内的语音数据。基于该生成数据集提取多维特征,并输入度量学习架构进行亲属关系验证。在包含年龄与亲属关系标签的KAN_AV语音数据集上开展实验,结果表明该方法显著提升了亲属关系验证的准确率,同时为未来研究提供了新思路。
原文摘要 · Abstract (English)
Audio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the task arises from differences in age across samples from different individuals, which can be interpreted as a domain bias in a cross-domain verification task. To address this issue, we design the notion of an "age-standardised domain" wherein we utilise the optimised CycleGAN-VC3 network to perform age-audio conversion to generate the in-domain audio. The generated audio dataset is employed to extract a range of features, which are then fed into a metric learning architecture to verify kinship. Experiments are conducted on the KAN_AV audio dataset, which contains age and kinship labels. The results demonstrate that the method markedly enhances the accuracy of kinship verification, while also offering novel insights for future kinship verification research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。