arXiv:2502.11544cs.CL2025-02被引 30

o1类大模型在多语言翻译中表现超群,尤其适合文化历史类翻译。

Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis

  • 通过分析推理机制,评估o1类模型在多语言翻译中的表现。
  • DeepSeek-R1在无上下文任务中超越GPT-4o,但中文输出易冗长。
  • 模型规模越大、温度越低,翻译越准确稳定,适合高精度场景。

o1类大模型通过模拟人类认知过程正重塑AI,但在多语言机器翻译(MMT)领域表现仍不明确。本研究评估了多个o1类大模型,并与ChatGPT、GPT-4o等传统模型对比。结果表明,o1类模型建立了新的多语言翻译基准:DeepSeek-R1在无上下文任务中超越GPT-4o;其在历史与文化翻译中表现优异,但中文输出存在冗余倾向。进一步分析揭示三点:(1) 高推理成本与慢速处理使复杂翻译任务资源开销更大;(2) 模型规模越大,常识推理与文化理解能力越强;(3) 温度参数显著影响输出质量——低温度带来更稳定、准确的翻译,高温度则降低连贯性与精确度。

原文摘要 · Abstract (English)

The o1-Like LLMs are transforming AI by simulating human cognitive processes, but their performance in multilingual machine translation (MMT) remains underexplored. This study examines: (1) how o1-Like LLMs perform in MMT tasks and (2) what factors influence their translation quality. We evaluate multiple o1-Like LLMs and compare them with traditional models like ChatGPT and GPT-4o. Results show that o1-Like LLMs establish new multilingual translation benchmarks, with DeepSeek-R1 surpassing GPT-4o in contextless tasks. They demonstrate strengths in historical and cultural translation but exhibit a tendency for rambling issues in Chinese-centric outputs. Further analysis reveals three key insights: (1) High inference costs and slower processing speeds make complex translation tasks more resource-intensive. (2) Translation quality improves with model size, enhancing commonsense reasoning and cultural translation. (3) The temperature parameter significantly impacts output quality-lower temperatures yield more stable and accurate translations, while higher temperatures reduce coherence and precision.

大模型翻译推理能力多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。