arXiv:2505.21954cs.CVcs.AI2025-05中稿 · Interspeech 2026被引 3

新数据集UniTalk挑战真实场景下的说话人检测,推动模型更鲁棒。

Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness

  • 构建涵盖多种真实复杂场景的新数据集UniTalk,强调泛化能力。
  • 顶尖模型在AVA上表现接近完美,但在UniTalk上性能大幅下降。
  • 用UniTalk训练的模型对真实视频更具泛化性,适合实际应用研究。

我们提出UniTalk,一个聚焦于挑战性场景的新型数据集,旨在提升说话人检测(ASD)模型的泛化能力。以往基准如AVA主要包含旧电影,与真实视频存在显著领域差距。相比之下,UniTalk涵盖多样视频类型,反映真实世界中的复杂条件,包括非主流语言、嘈杂背景和人群密集场景,同时在规模上与AVA相当。大量实验表明,在真实条件下ASD仍未解决:在AVA上表现近乎完美的模型在UniTalk上未达饱和。相反,使用UniTalk训练的模型在现代真实场景数据集Talkies和ASW上展现出更强的泛化能力。因此,UniTalk为ASD设立了新基准,为开发和评估通用且稳健的模型提供了宝贵资源。

原文摘要 · Abstract (English)

We present UniTalk, a novel dataset emphasizing challenging scenarios to enhance model generalization for the task of active speaker detection (ASD). Previously established benchmarks such as AVA predominantly comprise old movies and thus exhibit significant domain gaps with real-world video. In contrast, UniTalk covers diverse video types reflecting challenging real-world conditions, including underrepresented languages, noisy backgrounds, and crowded scenes, while being on par with AVA in scale. Extensive evaluations reveal that ASD remains unsolved under realistic conditions: state-of-the-art models near-perfect on AVA fail to reach saturation on UniTalk. Conversely, models trained on UniTalk generalize better to modern in-the-wild datasets including Talkies and ASW. UniTalk thus establishes a new benchmark for ASD, providing researchers with a valuable resource for developing and evaluating versatile and resilient models.

说话人检测真实场景数据集泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。