arXiv:2605.11993cs.CL2026-05

用视觉信息提升印地语系电影字幕翻译质量,效果更准且省计算。

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

论文配图:Towards Visually-Guided Movie Subtitle Translation for Indic Languages
图 1 · 摘自论文原文
  • 基于画面片段生成粗粒度属性摘要,捕捉情绪与上下文线索。
  • 仅替换20%-30%低质量段落,就能稳定提升翻译质量。
  • 适合资源少的印地语系语言,减少对视觉数据依赖。

电影字幕翻译本质上是多模态任务,但纯文本系统常忽略情感、动作和社会语境等视觉线索,尤其在低资源印地语系语言(英译印地语、孟加拉语、泰卢固语、泰米尔语和卡纳达语)中更为明显。我们针对五部完整电影进行了案例研究,对比了两种轻量级视觉定位策略:基于5分钟滑动窗口的结构化属性摘要,以及对字幕间视觉空白的自由文本摘要。分析表明,字幕与画面在时间上的错位是长视频中的主要障碍,导致泛化视觉定位效果不佳。然而,使用‘理想选择性定位’——仅将基线中最低质量的20%-30%段落替换为视觉增强输出——能持续提升COMET得分,且所需视觉处理量极低。在两种方法中,粗粒度属性摘要更具鲁棒性,能有效捕捉场景级情感和文本遗漏的细微语境。

原文摘要 · Abstract (English)

Movie subtitle translation is inherently multimodal, yet text-only systems often miss visual cues needed to convey emotion, action, and social nuance, especially for low-resource Indic languages (English to Hindi, Bengali, Telugu, Tamil and Kannada). We present a case study on five full-length films and compare two lightweight visual grounding strategies: structured attribute summaries from a 5-minute sliding window and free-text summaries of inter-subtitle visual gaps. Our analysis shows that temporal misalignment between subtitles and frames is a major obstacle in long-form video, often rendering indiscriminate visual grounding ineffective. However, oracle selective grounding, which replaces only the lowest-quality 20-30\% of baseline segments with visual-enhanced outputs, consistently improves COMET over the text-only baseline while requiring far less visual processing. Among the two approaches, coarse attribute-based visual context summarization is more robust, capturing scene-level emotion and contextual subtle cues that text alone often misses

字幕翻译多模态低资源语言视觉引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。