构建纪录片多模态翻译数据集,支持跨领域语义理解
TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries
- 按经济、自然等8个主题构建视频字幕对数据集
- 视觉信息提升翻译性能,全局上下文进一步改善结果
- 适合研究跨域迁移与上下文建模的多模态翻译者
现有多模态机器翻译(MMT)数据集多由静态图像或短片段构成,缺乏覆盖广泛领域和主题的长视频数据,难以满足纪录片翻译等真实场景需求。本文构建了TopicVD——一个基于主题的纪录片视频引导多模态机器翻译数据集,涵盖经济、自然等8个主题,包含视频-字幕对,并保留其上下文信息,以支持领域自适应与全局上下文建模研究。为更好捕捉图文共享语义,提出基于跨模态双向注意力模块的MMT模型。在TopicVD上的实验表明,视觉信息持续提升NMT模型性能;然而,模型在跨领域场景下性能显著下降,凸显领域自适应的重要性;同时,全局上下文能有效提升翻译质量。
原文摘要 · Abstract (English)
Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet the demands of real-world MMT tasks, such as documentary translation. In this study, we developed TopicVD, a topic-based dataset for video-supported multimodal machine translation of documentaries, aiming to advance research in this field. We collected video-subtitle pairs from documentaries and categorized them into eight topics, such as economy and nature, to facilitate research on domain adaptation in video-guided MMT. Additionally, we preserved their contextual information to support research on leveraging the global context of documentaries in video-guided MMT. To better capture the shared semantics between text and video, we propose an MMT model based on a cross-modal bidirectional attention module. Extensive experiments on the TopicVD dataset demonstrate that visual information consistently improves the performance of the NMT model in documentary translation. However, the MMT model's performance significantly declines in out-of-domain scenarios, highlighting the need for effective domain adaptation methods. Additionally, experiments demonstrate that global context can effectively improve translation performance. % Dataset and our implementations are available at https://github.com/JinzeLv/TopicVD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。