arXiv:2507.20752cs.CLcs.LG2025-07Conference of the …被引 1

无需人工标注,用合成数据训练多语言事实性评估模型。

Multilingual Self-Taught Faithfulness Evaluators

  • 仅用合成多语言摘要数据训练,不依赖人工标注。
  • 跨语言迁移学习使模型在多种语言上表现优于现有基线。
  • 适合需要低成本多语言评估的开发者和研究者。

大型语言模型(LLM)的广泛应用加剧了对自动评估系统的需求,尤其是应对信息幻觉问题。尽管现有事实性评估方法已展现潜力,但主要集中在英语,且通常需要昂贵的人工标注数据来微调专用模型。随着LLM在多语言场景中的普及,亟需无需大量标注数据即可跨语言运行的事实性评估工具。本文提出多语言自训练事实性评估框架,仅通过合成多语言摘要数据进行训练,并利用跨语言迁移学习。实验表明,大语言模型的通用语言能力与其在特定语言评估任务中的表现存在稳定关联。该框架在多个语言上均优于现有基线,包括最先进的英文评估器及基于机器翻译的方法。

原文摘要 · Abstract (English)

The growing use of large language models (LLMs) has increased the need for automatic evaluation systems, particularly to address the challenge of information hallucination. Although existing faithfulness evaluation approaches have shown promise, they are predominantly English-focused and often require expensive human-labeled training data for fine-tuning specialized models. As LLMs see increased adoption in multilingual contexts, there is a need for accurate faithfulness evaluators that can operate across languages without extensive labeled data. This paper presents Self-Taught Evaluators for Multilingual Faithfulness, a framework that learns exclusively from synthetic multilingual summarization data while leveraging cross-lingual transfer learning. Through experiments comparing language-specific and mixed-language fine-tuning approaches, we demonstrate a consistent relationship between an LLM's general language capabilities and its performance in language-specific evaluation tasks. Our framework shows improvements over existing baselines, including state-of-the-art English evaluators and machine translation-based approaches.

多语言事实性评估自训练大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。