简单管道比复杂系统更有效,专用于中世纪拉丁手稿翻译。
When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

- 用专用OCR+VLM的简单流水线,直接处理手稿图像。
- 专用OCR使字符错误率降低4.3倍,参数量却少得多。
- 适合低资源历史文本翻译,避免复杂模块带来的误差累积。
尽管机器翻译进展显著,视觉语言模型(VLM)在历史手稿上仍表现不佳,该领域对自然语言处理的核心能力提出挑战:低资源转写、古语词汇和噪声输入。我们提出一个系统框架,评估中世纪拉丁手稿从图像到翻译的全流程。在CATMuS拉丁语数据集上,领域专用光学字符识别(OCR)模型相比通用VLM,字符错误率降低达4.3倍,且参数量少多个数量级。我们引入全新数据集Interpres-Parallel-Corpus(IPC),包含1,383行对齐的手稿图像、转写与专家翻译,是首个针对中世纪拉丁语的数据集。实验揭示复杂性悖论:最简单的管道——专用OCR直接接入VLM——优于所有多组件变体。添加检索增强生成(RAG)或后OCR修正会引发提示饱和和错误传播,降低整体翻译质量。研究提供了新基准与低资源历史场景部署的实用指导。
原文摘要 · Abstract (English)
Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals. We present a systematic framework for evaluating the full image-to-translation pipeline on medieval Latin manuscripts, a setting in which scribal shorthand, ligatures, and parchment degradation expose failure modes that are invisible in clean-text benchmarks. Benchmarking on the CATMuS Latin dataset reveals a specialization gap: domain-specific Optical Character Recognition (OCR) models reduce character error rate by up to 4.3$\times$ compared to general-purpose VLMs, despite operating at orders of magnitude fewer parameters. We introduce the Interpres-Parallel-Corpus (IPC), a novel dataset comprising 1,383 aligned manuscript image lines, transcriptions, and expert translations, the first of its kind for medieval Latin. Our experiments uncover a complexity paradox: the simplest pipeline, a specialized OCR model feeding directly into a VLM, outperforms all multi-component variants. Adding retrieval-augmented generation (RAG) or post-OCR correction introduces prompt saturation and error propagation that degrade aggregate translation quality. These findings offer both a new benchmark and practical guidance for deploying translation systems in low-resource historical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。