arXiv:2504.17315cs.CVcs.AI2025-04

华为用大模型实现复杂版面文档图像端到端翻译

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

  • 用多任务学习+感知思维链训练大视觉语言模型
  • 统一框架处理有OCR和无OCR的文档翻译任务
  • 适合需要高精度文档翻译的工业场景

本文介绍了华为翻译服务中心(HW-TSC)在第19届国际文档分析与识别会议(DIMT25@ICDAR2025)“复杂版面文档图像端到端机器翻译”竞赛中的技术方案。基于先进的开源大视觉语言模型(LVLM),提出一种结合多任务学习与感知思维链的训练框架,构建了完整的端到端文档翻译系统。推理阶段采用最小贝叶斯解码与后处理策略,进一步提升翻译性能。该方案首次在统一框架下同时解决基于OCR与无OCR的文档图像翻译任务。本文系统阐述了训练方法、推理策略、基线模型、训练数据、实验设置及结果,验证了该方法在文档图像机器翻译中的有效性。

原文摘要 · Abstract (English)

This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" competition at the 19th International Conference on Document Analysis and Recognition (DIMT25@ICDAR2025). Leveraging state-of-the-art open-source large vision-language model (LVLM), we introduce a training framework that combines multi-task learning with perceptual chain-of-thought to develop a comprehensive end-to-end document translation system. During the inference phase, we apply minimum Bayesian decoding and post-processing strategies to further enhance the system's translation capabilities. Our solution uniquely addresses both OCR-based and OCR-free document image translation tasks within a unified framework. This paper systematically details the training methods, inference strategies, LVLM base models, training data, experimental setups, and results, demonstrating an effective approach to document image machine translation.

文档翻译视觉语言模型端到端多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。