首个科米-雅兹瓦-俄语平行语料库,助力濒危语言机器翻译评估
A Komi-Yazva--Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation

- 构建首个科米-雅兹瓦-俄语平行语料库,含457句对
- 检索式少样本提示比零样本更优,但提升有限
- 提供可复现的评估框架,适合濒危语言研究者
我们提出了首个科米-雅兹瓦-俄语平行语料库及明确的评估协议,用于研究濒危、极低资源语言环境下的大模型翻译。该数据集包含来自74篇叙事文本的457个对齐句对,附带文档化来源、句级对齐与故事标识符,支持防泄漏评估。基于此,我们在零样本和基于检索的少样本条件下对比现代大模型在科米-雅兹瓦到俄语翻译中的表现。评估协议包括故事级交叉验证、确定性检索、严格输出验证、参考与人工评分双指标,以及故事级不确定性估计。实验表明,大模型能生成非平凡翻译,但性能随模型族与提示方式差异显著;基于检索的少样本优于零样本,但额外检索上下文带来的增益有限。结果强调,该场景下的评估结论高度依赖指标选择与错误处理策略,因此本文将语料库定位为兼具数据贡献与可复现评估测试平台的价值。
原文摘要 · Abstract (English)
We present the first Komi-Yazva--Russian parallel corpus together with an explicit evaluation protocol for studying LLM translation in an endangered, extremely low-resource setting. The dataset contains 457 aligned sentence pairs from 74 narrative texts and is accompanied by documented provenance, sentence-level alignment, and story identifiers that enable leakage-aware evaluation. We use this setup to compare modern large language models on Komi-Yazva-to-Russian translation under severe parallel-data scarcity in zero-shot and retrieval-based few-shot regimes. The protocol includes story-level cross-validation, deterministic retrieval for few-shot prompting, strict validation of generated outputs, complementary reference-based and judge-based metrics, and story-level uncertainty estimates. Across models, LLMs produce non-trivial translations, but performance varies strongly by model family and prompting regime. Retrieval-based few-shot prompting consistently improves over zero-shot prompting, while gains beyond a small retrieved context remain limited. The results show that evaluative conclusions in this setting depend materially on metric choice and failure handling, so the paper frames the corpus as both a dataset contribution and a reproducible evaluation testbed for endangered-language machine translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。