o1类大模型在多语言翻译中表现超群,尤其适合文化历史类翻译。
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
- 通过分析推理机制,评估o1类模型在多语言翻译中的表现。
- DeepSeek-R1在无上下文任务中超越GPT-4o,但中文输出易冗长。
- 模型规模越大、温度越低,翻译越准确稳定,适合高精度场景。
o1类大模型通过模拟人类认知过程正重塑AI,但在多语言机器翻译(MMT)领域表现仍不明确。本研究评估了多个o1类大模型,并与ChatGPT、GPT-4o等传统模型对比。结果表明,o1类模型建立了新的多语言翻译基准:DeepSeek-R1在无上下文任务中超越GPT-4o;其在历史与文化翻译中表现优异,但中文输出存在冗余倾向。进一步分析揭示三点:(1) 高推理成本与慢速处理使复杂翻译任务资源开销更大;(2) 模型规模越大,常识推理与文化理解能力越强;(3) 温度参数显著影响输出质量——低温度带来更稳定、准确的翻译,高温度则降低连贯性与精确度。
原文摘要 · Abstract (English)
The o1-Like LLMs are transforming AI by simulating human cognitive processes, but their performance in multilingual machine translation (MMT) remains underexplored. This study examines: (1) how o1-Like LLMs perform in MMT tasks and (2) what factors influence their translation quality. We evaluate multiple o1-Like LLMs and compare them with traditional models like ChatGPT and GPT-4o. Results show that o1-Like LLMs establish new multilingual translation benchmarks, with DeepSeek-R1 surpassing GPT-4o in contextless tasks. They demonstrate strengths in historical and cultural translation but exhibit a tendency for rambling issues in Chinese-centric outputs. Further analysis reveals three key insights: (1) High inference costs and slower processing speeds make complex translation tasks more resource-intensive. (2) Translation quality improves with model size, enhancing commonsense reasoning and cultural translation. (3) The temperature parameter significantly impacts output quality-lower temperatures yield more stable and accurate translations, while higher temperatures reduce coherence and precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。