arXiv:2604.11283cs.CV2026-04综述

用多模态大模型统一视频翻译,提升语义与情感一致性

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

  • 按语义理解、表达生成、视觉合成三角色分类研究方法
  • 强调跨模态对齐与情感表达,实现更自然的视频翻译
  • 适合关注多模态生成与跨语言传播的研究者

多模态大语言模型(MLLMs)正推动视频翻译从语音识别、机器翻译、文本转语音和口型同步的串联流程,转向统一的多模态推理与生成问题。高质量视频翻译需兼顾语义保真度、时间对齐性、说话人一致性及情感表现力。本综述基于角色导向分类法,将相关研究划分为三类功能角色:语义推理者(融合视频理解、时间推理与多模态融合)、情感表演者(支持可控且上下文感知的语音生成)、视觉合成者(实现口型同步与视觉连贯的说话人呈现)。我们总结各角色代表性数据集、基准测试与评估指标,并指出当前评测体系难以满足端到端视频翻译需求。最后,提出长时视频理解、时间建模、多模态对齐、多语言鲁棒性及负责任部署等开放挑战,展望自然可信的跨语言视频交流未来方向。

原文摘要 · Abstract (English)

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a unified multimodal reasoning and generation problem. High-quality video translation requires not only semantic fidelity, but also temporal alignment, speaker consistency, and emotional expressiveness across visual, acoustic, and linguistic streams. This survey provides a focused review of MLLM-enabled video translation through a role-oriented taxonomy. We organize MLLM-enabled and MLLM-relevant studies into three functional roles: Semantic Reasoner, which grounds translation in video understanding, temporal reasoning, and multimodal fusion; Expressive Performer, which supports controllable and context-aware speech generation; and Visual Synthesizer, which enables lip synchronization and visually coherent speaker rendering. We further summarize representative datasets, benchmarks, and metrics for each role, and discuss how current evaluation protocols fall short of end-to-end video translation requirements. Finally, we identify open challenges in long-form video understanding, temporal modeling, multimodal alignment, multilingual robustness, and responsible deployment, outlining future directions for natural and trustworthy cross-lingual video communication.

视频翻译多模态大模型跨语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。