arXiv:2511.09374cs.CLcs.AI2025-11中稿 · IJCNLP-AACL 2025 F…被引 2

多语言文本质量评估新框架,让大模型自动判断文本好坏

MTQ-Eval: Multilingual Text Quality Evaluation for Language Models

  • 用高质量和低质量文本训练模型,学习文本质量判断
  • 覆盖115种语言,评估效果显著提升
  • 可提升下游任务表现,适合多语言应用开发者

大语言模型用于输出评估正变得越来越高效且可扩展。然而,这种能力是否能从特定任务延伸到更通用的文本质量评估,尤其是在多语言环境下,仍不明确。本文提出 MTQ-Eval,一种新型多语言文本质量评估框架,通过学习高质量与低质量文本样例来调整内部表示。我们首先自动生成文本质量偏好数据,再用其训练开源基础大模型,使其对齐高/低质量文本的评分。在115种语言上的综合评估显示,该模型性能显著提升。进一步分析表明,这种增强的评估能力也带来了下游任务的明显改进。

原文摘要 · Abstract (English)

The use of large language models (LLMs) for evaluating outputs is becoming an increasingly effective and scalable approach. However, it remains uncertain whether this capability extends beyond task-specific evaluations to more general assessments of text quality, particularly in multilingual contexts. In this study, we introduce, MTQ-Eval, a novel framework for multilingual text quality evaluation that learns from examples of both high- and low-quality texts, adjusting its internal representations. To develop MTQ-Eval, we first automatically generate text quality preference data and then use it to train open-source base LLMs to align with ratings of high- and low-quality text. Our comprehensive evaluation across 115 languages demonstrates the improved performance of the proposed model. Upon further analysis, we find that this enhanced evaluation capability also leads to notable improvements in downstream tasks.

文本评估多语言LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。