arXiv:2512.16287cs.CL2025-12被引 2

对比GPT推理与非推理模型在濒危乌拉尔语翻译中的表现

Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures

  • 用文学文本并行语料库测试GPT模型翻译意愿
  • 推理模型拒绝率比非推理模型低16个百分点
  • 对濒危语言保护研究者有重要参考价值

大型语言模型(LLMs)在翻译任务中的评估主要集中于高资源语言,对低资源及濒危语言的表现仍缺乏了解。本研究系统比较了OpenAI GPT系列模型在芬兰语与四种低资源乌拉尔语(科米-兹里亚尼语、莫克沙语、埃爾齊亞語、烏德穆爾特語)之间翻译的表现,重点分析推理型与非推理型架构的差异。基于文学文本并行语料库,通过拒绝率分析评估模型的翻译意愿。结果表明,推理型模型拒绝翻译的比例比非推理型低16个百分点,显示出更强的翻译主动性。研究为乌拉尔语研究者和濒危语言保护实践提供了重要参考,深化了对推理模型在语言保存中作用的理解。

原文摘要 · Abstract (English)

The evaluation of Large Language Models (LLMs) for translation tasks has primarily focused on high-resource languages, leaving a significant gap in understanding their performance on low-resource and endangered languages. This study presents a comprehensive comparison of OpenAI's GPT models, specifically examining the differences between reasoning and non-reasoning architectures for translating between Finnish and four low-resource Uralic languages: Komi-Zyrian, Moksha, Erzya, and Udmurt. Using a parallel corpus of literary texts, we evaluate model willingness to attempt translation through refusal rate analysis across different model architectures. Our findings reveal significant performance variations between reasoning and non-reasoning models, with reasoning models showing 16 percentage points lower refusal rates. The results provide valuable insights for researchers and practitioners working with Uralic languages and contribute to the broader understanding of reasoning model capabilities for endangered language preservation.

语言保护GPT低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。