新基准揭示多语言大模型翻译幻觉的4类触发机制
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
- 按指令与原文偏离分离构建诊断分类法
- 覆盖11种语言方向,含5435条人工验证数据
- 可识别模型规模、长度敏感等幻觉模式
大语言模型在机器翻译领域取得进展,但仍易产生幻觉。现有翻译基准难以暴露多语言LLM的缺陷。为此,本文提出一种诊断框架,基于将‘指令偏离’与‘源文偏离’分离的分类法,构建了跨11个英译其他语言方向的多语言人工验证基准HalloMTBench。通过4个前沿大模型生成候选译文,并经由多模型评委组与专家联合验证,最终筛选出5435条高质量样本。在该基准上评估了17个大模型,发现存在四类独特的‘幻觉触发因素’:模型规模差异、源文本长度敏感性、语言偏见以及强化学习放大的语言混杂现象。HalloMTBench为诊断大模型翻译失败提供了前瞻性测试平台。数据集已开源于https://huggingface.co/collections/AIDC-AI/marco-mt。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have advanced machine translation but remain vulnerable to hallucinations. Unfortunately, existing MT benchmarks are not capable of exposing failures in multilingual LLMs. To disclose hallucination in multilingual LLMs, we introduce a diagnostic framework with a taxonomy that separates Instruction Detachment from Source Detachment. Guided by this taxonomy, we create HalloMTBench, a multilingual, human-verified benchmark across 11 English-to-X directions. We employed 4 frontier LLMs to generate candidates and scrutinize these candidates with an ensemble of LLM judges, and expert validation. In this way, we curate 5,435 high-quality instances. We have evaluated 17 LLMs on HalloMTBench. Results reveal distinct ``hallucination triggers'' -- unique failure patterns reflecting model scale, source length sensitivity, linguistic biases, and Reinforcement-Learning (RL) amplified language mixing. HalloMTBench offers a forward-looking testbed for diagnosing LLM translation failures. HalloMTBench is available in https://huggingface.co/collections/AIDC-AI/marco-mt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。