分析10万次翻译请求,发现学生用手机多向蒂顿语翻译科普医疗内容。
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
- 通过观察tetun.org实际使用数据,发现用户主要将高资源语言译为蒂顿语。
- 80%请求来自移动设备,覆盖教育、医疗等日常领域,非新闻类。
- 建议低资源语言系统优先提升教育场景下的翻译准确率。
低资源机器翻译(MT)面临多样化的社区需求与应用挑战,但现有认知仍不充分。为补充依赖小样本的问卷与焦点小组研究,本文对蒂顿语专用MT服务tetun.org的实际使用模式开展观察性研究。分析10万条翻译请求发现,用户(多为移动设备上的学生)普遍将高资源语言文本翻译为蒂顿语,涵盖科学、医疗及日常生活等多样化领域,与现有蒂顿语语料库以政府和社会议题新闻为主的特征形成鲜明对比。结果表明,面向如蒂顿语这类制度化少数语言的MT系统应优先关注教育相关领域的准确性,且采用高资源到低资源方向的翻译。本研究展示了观察分析如何基于真实社区需求,推动低资源语言技术的发展。
原文摘要 · Abstract (English)
Low-resource machine translation (MT) presents a diversity of community needs and application challenges that remain poorly understood. To complement surveys and focus groups, which tend to rely on small samples of respondents, we propose an observational study on actual usage patterns of tetun$.$org, a specialized MT service for the Tetun language, which is the lingua franca in Timor-Leste. Our analysis of 100,000 translation requests reveals patterns that challenge assumptions based on existing corpora. We find that users, many of them students on mobile devices, typically translate text from a high-resource language into Tetun across diverse domains including science, healthcare, and daily life. This contrasts sharply with available Tetun corpora, which are dominated by news articles covering government and social issues. Our results suggest that MT systems for institutionalized minority languages like Tetun should prioritize accuracy on domains relevant to educational contexts, in the high-resource to low-resource direction. More broadly, this study demonstrates how observational analysis can inform low-resource language technology development, by grounding research in practical community needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。