arXiv:2505.19714cs.CLcs.AI2025-05被引 5

用多任务强化学习让图文翻译模型一次搞定识别、理解与翻译。

MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning

  • 设计多任务强化学习框架,统一优化文本识别、上下文推理和翻译能力。
  • 在MIT-10M上超越Qwen2.5-VL-72B等大模型,多个指标领先。
  • 首个社交媒体图文翻译评测集XHSPost,适合真实场景评估。

文本图像机器翻译(TIMT)——即翻译嵌入图像中的文字内容——对无障碍访问、跨语言信息获取和现实文档理解至关重要。然而,由于需要高精度光学字符识别(OCR)、鲁棒的视觉-文本推理和高质量翻译,该任务仍具挑战性,常依赖级联的多阶段流水线。近期大规模强化学习(RL)提升了大语言模型(LLMs)和多模态大模型(MLLMs)的推理能力,但其在端到端TIMT中的应用仍不充分。为此,我们提出MT³,首个将多任务强化学习应用于MLLM实现端到端TIMT的框架。MT³采用多任务优化范式,聚焦文本识别、上下文感知推理与翻译三项核心子技能,通过新颖的多混合奖励机制,将规则化强化学习策略适配于TIMT的复杂性,提供细粒度、非二值化的跨任务反馈。此外,为促进真实跨文化社交媒体场景下的TIMT评估,我们引入首个社交媒体TIMT基准XHSPost。MT³-7B-Zero在最新的域内MIT-10M基准上取得当前最优结果,显著优于Qwen2.5-VL-72B和InternVL2.5-78B等强基线。模型还展现出对分布外语言对和数据集的良好泛化能力。深入分析揭示了多任务协同、强化学习初始化、课程设计及奖励制定对提升MLLM驱动的TIMT的关键作用。

原文摘要 · Abstract (English)

Text Image Machine Translation (TIMT)-the task of translating textual content embedded in images-is critical for applications in accessibility, cross-lingual information access, and real-world document understanding. However, TIMT remains a complex challenge due to the need for accurate optical character recognition (OCR), robust visual-text reasoning, and high-quality translation, often requiring cascading multi-stage pipelines. Recent advances in large-scale Reinforcement Learning (RL) have improved reasoning in Large Language Models (LLMs) and Multimodal LLMs (MLLMs), but their application to end-to-end TIMT is still underexplored. To bridge this gap, we introduce MT$^{3}$, the first framework to apply Multi-Task RL to MLLMs for end-to-end TIMT. MT$^{3}$ adopts a multi-task optimization paradigm targeting three key sub-skills: text recognition, context-aware reasoning, and translation. It is trained using a novel multi-mixed reward mechanism that adapts rule-based RL strategies to TIMT's intricacies, offering fine-grained, non-binary feedback across tasks. Furthermore, to facilitate the evaluation of TIMT in authentic cross-cultural and real-world social media contexts, we introduced XHSPost, the first social media TIMT benchmark. Our MT$^{3}$-7B-Zero achieves state-of-the-art results on the latest in-domain MIT-10M benchmark, outperforming strong baselines such as Qwen2.5-VL-72B and InternVL2.5-78B by notable margins across multiple metrics. Additionally, the model shows strong generalization to out-of-distribution language pairs and datasets. In-depth analyses reveal how multi-task synergy, reinforcement learning initialization, curriculum design, and reward formulation contribute to advancing MLLM-driven TIMT.

图文翻译多模态强化学习MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。