用提示工程提升大模型翻译编程题能力,让罗马尼亚题准确转英文。
Exploring Large Language Models for Translating Romanian Computational Problems into English
- 设计结构化提示,让大模型更好理解并翻译编程竞赛题。
- GPT-4o等模型在重复测试中表现稳定,准确率超人工参考。
- 适合需要多语言编程题自动翻译的教育和竞赛场景。
近期研究发现,大语言模型(LLMs)在将罗马尼亚语编程问题翻译成英语时,性能低于原生罗马尼亚语形式。准确翻译对编程竞赛自动翻译、高质量教育材料生成及防止人工翻译错误至关重要。本研究显示,通过优化提示设计,大模型可在翻译小众语言时保持甚至提升性能。我们评估了OpenRoLLM、Llama 3.1 8B、Llama 3.2 3B和GPT-4o等多种模型,在多次运行中验证其翻译准确性和稳定性。同时,我们对罗马尼亚县级信息学奥赛(OJI)数据集进行了高质量英文标注,增强其未来训练与评测价值。通过语法与语义分析,确认在人工监督下,大模型可成为多语言问题求解的可行方案。经认证专家对比评估,部分模型翻译质量已接近或超过人工水平,展现其在真实场景中的潜力。
原文摘要 · Abstract (English)
Recent studies have suggested that large language models (LLMs) underperform on mathematical and computer science tasks when these problems are translated from Romanian into English, compared to their original Romanian format. Accurate translation is critical for applications ranging from automatic translations in programming competitions to the creation of high-quality educational materials, as well as minimizing errors or fraud in human translations. This study shows that robust large language models (LLMs) can maintain or even enhance their performance in translating less common languages when given well-structured prompts. Our findings suggest that LLMs, with appropriate supervision, can be reliably used for the automatic translation of IOI (International Olympiad in Informatics)-style tasks. We evaluate several translation methods across multiple LLMs, including OpenRoLLM, Llama 3.1 8B, Llama 3.2 3B and GPT-4o, assessing their translation accuracy and performance stability through repeated runs. Additionally, we augment the OJI (Romanian County-Level Informatics Olympiad) Romanian dataset with accurate English translations, enhancing its utility for future LLM training and evaluation. Through detailed syntactic and semantic analyses, we confirm that with human oversight, LLMs can serve as a viable solution for multilingual problem-solving. We also compare the translation quality of LLMs against human translators, as evaluated by a certified expert, underscoring the potential of LLMs in realworld scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。