在真实音频档案中实现可靠语音检索,解决标签少与环境嘈杂难题。
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
- 基于有限元数据构建语音检索系统,结合语音分割与嵌入提取。
- 在干净和带噪环境下均表现稳定,验证系统鲁棒性。
- 适用于大规模老旧音视频档案,适合媒体存档与数字重建场景。
日益丰富的公开或企业音视频档案库凸显了高效内容检索的重要性。本文研究了‘真实世界’中语音检索系统的挑战、解决方案、有效性与鲁棒性,核心问题包括:从有限元数据中提取任务相关标签,以及应对档案中从安静演播室到复杂噪声环境的非约束声学条件。以1948至1979年公开的BBC Rewind档案为例,本框架可推广至大规模、可能老化且无内容控制的音视频档案。这些档案通常仅有简略通用描述,难以支持如语音检索等特定应用,而人工标注不可行。我们探讨了系统构建中的多个环节(如说话人分离、嵌入提取、查询选择),分析其挑战与可行性。通过在清洁环境及多种失真条件下系统性实验评估性能,结果表明所提系统兼具有效性和鲁棒性,证明该框架在广泛应用场景中的通用性与可扩展性。
原文摘要 · Abstract (English)
There is a growing abundance of publicly available or company-owned audio/video archives, highlighting the increasing importance of efficient access to desired content and information retrieval from these archives. This paper investigates the challenges, solutions, effectiveness, and robustness of speaker retrieval systems developed "in the wild" which involves addressing two primary challenges: extraction of task-relevant labels from limited metadata for system development and evaluation, as well as the unconstrained acoustic conditions encountered in the archive, ranging from quiet studios to adverse noisy environments. While we focus on the publicly-available BBC Rewind archive (spanning 1948 to 1979), our framework addresses the broader issue of speaker retrieval on extensive and possibly aged archives with no control over the content and acoustic conditions. Typically, these archives offer a brief and general file description, mostly inadequate for specific applications like speaker retrieval, and manual annotation of such large-scale archives is unfeasible. We explore various aspects of system development (e.g., speaker diarisation, embedding extraction, query selection) and analyse the challenges, possible solutions, and their functionality. To evaluate the performance, we conduct systematic experiments in both clean setup and against various distortions simulating real-world applications. Our findings demonstrate the effectiveness and robustness of the developed speaker retrieval systems, establishing the versatility and scalability of the proposed framework for a wide range of applications beyond the BBC Rewind corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。