用少量资源提升濒危语言翻译质量,发现平行语料最有效。
Testing the Limits of Machine Translation from One Book
- 用语法、词典和双语句子组合测试大模型翻译能力
- 平行语料在人工评估中显著优于其他方法
- 语法单独使用效果差,需结合语料才能提升准确率
当前先进模型可通过上下文学习实现对未见语言的翻译。本文聚焦于卡努里语——一种虽有大量使用者但数字资源极少的语言。我们构建了两个评估数据集:一个涵盖健康与人道主义术语,另一个为通用术语。通过提供语法、词典和双语句子的不同组合,评估大模型在翻译中的表现,并与母语者及语言学家的结果对比。采用自动指标与母语者对流畅性与准确性评估。结果显示,双语句子仍是最重要的数据来源,在人工评价与自动指标中均表现最佳;语法虽能改善零样本翻译,但无法独立支撑有效翻译;人类评估表明,大模型在语义准确度上优于流畅性。研究提示,评估大模型翻译应采用多维度标准,且仅靠语法不足以支持领域特定翻译。
原文摘要 · Abstract (English)
Current state-of-the-art models demonstrate capacity to leverage in-context learning to translate into previously unseen language contexts. Tanzer et al. [2024] utilize language materials (e.g. a grammar) to improve translation quality for Kalamang using large language models (LLMs). We focus on Kanuri, a language that, despite having substantial speaker population, has minimal digital resources. We design two datasets for evaluation: one focused on health and humanitarian terms, and another containing generalized terminology, investigating how domain-specific tasks impact LLM translation quality. By providing different combinations of language resources (grammar, dictionary, and parallel sentences), we measure LLM translation effectiveness, comparing results to native speaker translations and human linguist performance. We evaluate using both automatic metrics and native speaker assessments of fluency and accuracy. Results demonstrate that parallel sentences remain the most effective data source, outperforming other methods in human evaluations and automatic metrics. While incorporating grammar improves over zero-shot translation, it fails as an effective standalone data source. Human evaluations reveal that LLMs achieve accuracy (meaning) more effectively than fluency (grammaticality). These findings suggest LLM translation evaluation benefits from multidimensional assessment beyond simple accuracy metrics, and that grammar alone, without parallel sentences, does not provide sufficient context for effective domain-specific translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。