用多模态强化学习提升汉喃古籍到现代越南语的翻译质量。
Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment

- 结合图像与文本的多模态框架,通过偏好对齐优化生成。
- DPO方法在多个指标上超越基线,提升语义与词汇准确率。
- 适合历史文献数字化、低资源语言翻译研究者使用。
将汉喃古籍翻译为现代越南语面临图像退化、生僻字多及平行语料稀缺等挑战。本文提出一种多模态强化学习偏好对齐框架,以古籍图像和对齐的汉喃原文为条件生成越南语。模型融合四种编码器:CLIP ViT-L/14@336(视觉特征)、bert-base-chinese(汉喃表示)、vinai/phobert-base(越南语表示)和T5-small编码器状态。各模态特征经特定投影后,通过融合模块压缩为512维共享表示。在相同微调策略基础上,对比PPO、DPO和KTO三种方法在工作级宏平均评估下的表现。DPO在BLEU-4、ROUGE-L、BERTScore、语义相似度、字符错误率(CER)、词错误率(WER)和词元准确率上最优;PPO在精确率、召回率和F1值上最高;KTO凭借其理想-非理想效用目标保持竞争力。所有偏好对齐策略均显著提升基线监督微调模型的BLEU-4与语义相似度得分。结果表明,多模态偏好优化能有效弥补监督学习不足,提升低资源历史翻译中的词汇与语义质量。
原文摘要 · Abstract (English)
Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF preference-alignment framework that conditions Vietnamese generation on manuscript images and aligned Han-Nom source text. The model combines four streams: CLIP ViT-L/14@336 for visual features, bert-base-chinese for Han-Nom representations, vinai/phobert-base for Vietnamese representations, and T5-small encoder states. Modality-specific projections and a fusion block compress the resulting 2,048-dimensional concatenation into a shared 512-dimensional representation. Starting from the same supervised fine-tuned policy, we compare PPO, DPO, and KTO under matched work-level macro-averaged evaluation. DPO achieves the best BLEU-4, ROUGE-L, BERTScore, semantic similarity, CER, WER, and token accuracy, whereas PPO obtains the highest precision, recall, and F1. KTO remains competitive through its desirable-undesirable utility objective. All preference-aligned policies improve the BLEU-4 and semantic-similarity scores available for the SFT baseline. These results indicate that multimodal preference optimization complements supervised learning by improving lexical and semantic quality in low-resource historical translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。