arXiv:2410.11068cs.CVcs.LG2024-10被引 3

融合音视频线索,精准识别影视剧对话中谁在说话

Character-aware audio-visual subtitling in context

  • 结合语音、音频与视觉信息,定位说话人脸并匹配角色身份
  • 通过上下文推理,将短对话段落准确归属到对应角色
  • 适用于影视剧字幕生成,提升角色识别与说话人分离精度

本文提出一种改进的电视连续剧场景下角色感知的音视频字幕生成框架。该方法整合语音识别、说话人分离与角色识别技术,同时利用音频和视觉线索,解决‘说了什么’、‘何时说’以及‘谁在说’三个核心问题。首先,利用音视频同步性从多个人物画面中识别出正在说话的脸,并为对应语音段分配角色身份,显著提升识别准确率。其次,针对短时对话片段难以判定说话人的问题,提出基于局部语音嵌入与大语言模型文本推理的方法,借助场景内对话上下文实现更精准的说话人归属。在包含12部电视剧的数据集上验证,本方法在说话人分离与角色识别任务上均优于现有方法。

原文摘要 · Abstract (English)

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This holistic solution addresses what is said, when it's said, and who is speaking, providing a more comprehensive and accurate character-aware subtitling for TV shows. Our approach brings improvements on two fronts: first, we show that audio-visual synchronisation can be used to pick out the talking face amongst others present in a video clip, and assign an identity to the corresponding speech segment. This audio-visual approach improves recognition accuracy and yield over current methods. Second, we show that the speaker of short segments can be determined by using the temporal context of the dialogue within a scene. We propose an approach using local voice embeddings of the audio, and large language model reasoning on the text transcription. This overcomes a limitation of existing methods that they are unable to accurately assign speakers to short temporal segments. We validate the method on a dataset with 12 TV shows, demonstrating superior performance in speaker diarisation and character recognition accuracy compared to existing approaches. Project page : https://www.robots.ox.ac.uk/~vgg/research/llr-context/

字幕生成音视频对齐角色识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。