用全局视频上下文提升长视频字幕翻译效果
Video-guided Machine Translation with Global Video Context
- 通过语义编码与向量库检索构建相关视频片段上下文集
- 在长视频上显著优于基线模型,翻译更连贯
- 适合需要理解跨段落叙事的多模态翻译任务
近年来,视频引导的多模态翻译(VMT)取得显著进展。然而,现有方法大多依赖局部对齐的视频片段与字幕一一对应,难以捕捉长视频中跨多个片段的全局叙事上下文。为此,我们提出一种全局视频引导的多模态翻译框架,利用预训练语义编码器和基于向量数据库的字幕检索,构建与目标字幕语义密切相关的视频片段上下文集。采用注意力机制聚焦高相关视觉内容,同时保留其余视频特征以维持更广泛上下文信息。此外,设计区域感知的跨模态注意力机制,增强翻译过程中的语义对齐。在大规模纪录片翻译数据集上的实验表明,该方法显著优于基线模型,尤其在长视频场景下表现突出。
原文摘要 · Abstract (English)
Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative context across multiple segments in long videos. To overcome this limitation, we propose a globally video-guided multimodal translation framework that leverages a pretrained semantic encoder and vector database-based subtitle retrieval to construct a context set of video segments closely related to the target subtitle semantics. An attention mechanism is employed to focus on highly relevant visual content, while preserving the remaining video features to retain broader contextual information. Furthermore, we design a region-aware cross-modal attention mechanism to enhance semantic alignment during translation. Experiments on a large-scale documentary translation dataset demonstrate that our method significantly outperforms baseline models, highlighting its effectiveness in long-video scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。