借鉴机器翻译评估经验,提升多语言大模型生成能力的评测科学性。
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
- 从机器翻译评估中汲取经验,优化多语言模型生成评测流程。
- 实验证明新方法能更准确揭示不同模型间的质量差异。
- 提供可操作的评测检查清单,适合研究者与开发者使用。
多语言大语言模型(mLLMs)的生成能力与语言覆盖范围快速进步,但其生成能力的评估仍缺乏全面性、科学性及跨实验室的一致性,制约了对mLLMs发展的有效引导。本文借鉴机器翻译(MT)评估领域的发展经验,该领域曾面临类似挑战,并经过数十年积累形成了透明的报告标准与可靠的评估体系。通过在生成评估流程的关键阶段开展针对性实验,我们证明了借鉴MT评估最佳实践能更深入理解模型间质量差异。同时,我们识别出保障评估方法自身可靠性的核心要素,提出一套可用于稳健元评估的组件。最终,我们将这些洞见提炼为一套可操作的建议清单,供mLLM研究与开发参考。
原文摘要 · Abstract (English)
Generation capabilities and language coverage of multilingual large language models (mLLMs) are advancing rapidly. However, evaluation practices for generative abilities of mLLMs are still lacking comprehensiveness, scientific rigor, and consistent adoption across research labs, which undermines their potential to meaningfully guide mLLM development. We draw parallels with machine translation (MT) evaluation, a field that faced similar challenges and has, over decades, developed transparent reporting standards and reliable evaluations for multilingual generative models. Through targeted experiments across key stages of the generative evaluation pipeline, we demonstrate how best practices from MT evaluation can deepen the understanding of quality differences between models. Additionally, we identify essential components for robust meta-evaluation of mLLMs, ensuring the evaluation methods themselves are rigorously assessed. We distill these insights into a checklist of actionable recommendations for mLLM research and development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。