arXiv:2509.07829cs.CLcs.AI2025-09被引 2

用开源模型构建英罗文学翻译资源,提升低资源语言表现。

Building Large-Scale English-Romanian Literary Translation Resources with Open Models

  • 基于合成童话数据集,用大模型生成高质量罗马尼亚语译文。
  • 120亿参数模型经两阶段微调,在流畅性和准确性上接近商业模型。
  • 开源全流程工具链,适合研究低成本跨语言叙事生成的学者。

文学翻译作为机器翻译中的独立复杂任务近年受到关注,但小规模开源模型在低资源语言如罗马尼亚语上的表现仍不理想。本文提出TinyFabulist Translation Framework (TF2),一个统一的英文→罗马尼亚语文学翻译数据集构建、微调与评估框架。基于目前最大的合成英语寓言数据集DS-TF1-EN-3M,我们利用高性能大语言模型(LLM)生成了15,000条高质量罗马尼亚语参考译文。随后对120亿参数的开源模型进行两阶段微调:(i) 指令微调以捕捉特定文类叙事风格,(ii) 适配器压缩以实现高效部署。评估采用五维度大模型评分体系(准确性、流畅性、连贯性、风格、文化适配性)作为主评价框架,并辅以基于参考的句级BLEU分数。结果表明,微调后的模型TF2-12B在自动与人工评估中均表现出强流畅性与充分性,显著缩小了与顶级专有模型的差距,同时具备开放性、可访问性与显著更低的成本。我们公开发布微调模型及两个大规模合成平行语料库(DS-TF2-EN-RO-3M 和 DS-TF2-EN-RO-15K),以及所有脚本与评估提示。TF2为低成本翻译、跨语言叙事生成及低资源语种文化内容的开源模型应用提供了端到端可复现的解决方案。

原文摘要 · Abstract (English)

Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, particularly for low-resource languages such as Romanian. We introduce the TinyFabulist Translation Framework (TF2), a unified framework for dataset creation, fine-tuning, and evaluation in English $\to$ Romanian literary translation. Building on DS-TF1-EN-3M, the largest collection of synthetic English fables to date, our pipeline first generates 15k high-quality Romanian references from the TF1 pool using a high-performing large language model (LLM). We then apply a two-stage fine-tuning process to a 12B-parameter open-weight model: (i) instruction tuning to capture genre-specific narrative style, and (ii) adapter compression for efficient deployment. Evaluation combines a five-dimension LLM-based rubric (accuracy, fluency, coherence, style, cultural adaptation) as the primary comparative framework, alongside corpus-level Bilingual Evaluation Understudy (BLEU) reported as a secondary reference-based consistency metric. Our fine-tuned model (TF2-12B) achieves strong fluency and adequacy, narrowing the gap to top-performing proprietary models under automated and human-anchored evaluation, while being open, accessible, and significantly more cost-effective. We publicly release the fine-tuned model and two large-scale synthetic parallel datasets (DS-TF2-EN-RO-3M and DS-TF2-EN-RO-15K), along with all scripts and evaluation prompts. TF2 provides an end-to-end, reproducible pipeline for research on cost-efficient translation, cross-lingual narrative generation, and the broad adoption of open models for culturally significant literary content in low-resource settings.

文学翻译开源模型低资源语言叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。