arXiv:2603.16966cs.CVcs.AI2026-03中稿 · CVPR

融合视觉语音与字幕信息,实现影视类长视频的精准说话人分割。

CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

  • 利用视觉锚点聚类与语音语言模型协同定位说话人。
  • 在中英双语影视数据集上达到领先性能,支持大量说话人识别。
  • 适合影视内容分析、智能字幕生成等开放域应用场景。

传统说话人分离系统多聚焦于会议、访谈等受限场景,说话人数量有限且声学条件良好。为探索开放世界说话人分离,本文将其拓展至影视等复杂音视频内容领域,面临长视频理解、大量说话人、跨模态不同步及真实环境波动等挑战。为此,我们提出电影级说话人注册与分离框架(CineSRD),统一融合视频、语音与字幕中的视觉、听觉和语言线索进行说话人标注。CineSRD首先通过视觉锚点聚类完成初始说话人注册,再结合语音语言模型进行说话人片段检测,精炼标注并补充未现身说话人。此外,我们构建并发布了首个面向影视媒体的说话人分离基准数据集,涵盖中英文节目。实验表明,CineSRD在自建基准上表现优异,在传统数据集上也具竞争力,验证了其在开放世界音视频场景下的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.

说话人分离多模态影视分析开放域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。