构建古泰卢固语诗转现代文数据集,评测大模型翻译效果。
Translating Classical Poetry into Modern Prose
- 构建600首古典诗及其双语译文数据集。
- 大语言模型表现优于传统机器翻译系统。
- 发现中英文译文生成与评估存在系统性问题。
我们提出了Padyam2Gadyam,一个将13至17世纪泰卢固古典诗歌翻译为现代泰卢固语和英语散文的任务数据集。该数据集包含600首诗歌及其经人工验证的泰卢固语和英语散文译文。我们使用该数据集评估了两种机器翻译系统和五种当代大型语言模型在泰卢固语及英语诗歌转散文任务中的表现。结果表明,通用大语言模型在该任务上优于传统机器翻译系统,但各类系统在两种语言的散文生成与评估方面均存在系统性问题。
原文摘要 · Abstract (English)
We introduce Padyam2Gadyam a dataset for the task of poem-to-prose translation from 13th-17th Century Telugu Classical Poetry to contemporary Telugu and English prose. The dataset consists of 600 poems and their human-verified Telugu and English prose translations. We evaluated 2 machine translation systems and 5 contemporary Large Language Models (LLMs) on their ability to do poem-to-prose translation into Telugu and English using this dataset. Our results indicate that while the general purpose LLMs are better than the machine translation systems for this task, there are systematic issues with the generation and evaluation of prose translation in both languages across systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。