首个九语言文本净化评估基准,提升多语种风格迁移评测可靠性
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
- 构建九语言文本净化评估基准,涵盖阿拉伯语、中文、英语等
- 基于LLM的评估方法相比传统指标与人工判断相关性显著提升
- 为多语种文本净化提供可落地的评估指南,适合相关研究者使用
尽管大语言模型(LLMs)取得显著进展,但文本生成任务如文本风格迁移(TST)的可靠评估仍是开放挑战。现有研究显示,自动评估指标常与人类判断相关性较差(Dementieva et al., 2024;Pauli et al., 2025),限制了对模型性能的准确评估。此外,多数前期工作集中于英语,而多语言TST系统(尤其是文本净化)的评估仍严重不足。本文首次开展跨九种语言(阿拉伯语、阿姆哈拉语、中文、英语、德语、印地语、俄语、西班牙语、乌克兰语)的文本净化评估基准研究。受机器翻译评估启发,我们比较神经基自动指标与基于LLM的评估方法,并结合特定任务微调模型进行实验。分析表明,所提方法在与人工判断的相关性上显著优于基线。我们还提供了构建稳健、可靠的多语种文本净化评估流程的实际建议。
原文摘要 · Abstract (English)
Despite notable advances in large language models (LLMs), reliable evaluation of text generation tasks such as text style transfer (TST) remains an open challenge. Existing research has shown that automatic metrics often correlate poorly with human judgments (Dementieva et al., 2024; Pauli et al., 2025), limiting our ability to assess model performance accurately. Furthermore, most prior work has focused primarily on English, while the evaluation of multilingual TST systems, particularly for text detoxification, remains largely underexplored. In this paper, we present the first comprehensive multilingual benchmarking study of evaluation metrics for text detoxification evaluation across nine languages: Arabic, Amharic, Chinese, English, German, Hindi, Russian, Spanish, and Ukrainian. Drawing inspiration from machine translation evaluation, we compare neural-based automatic metrics with LLM-as-a-judge approaches together with experiments on task-specific fine-tuned models. Our analysis reveals that the proposed metrics achieve significantly higher correlation with human judgments compared to baseline approaches. We also provide actionable insights and practical guidelines for building robust and reliable multilingual evaluation pipelines for text detoxification and related TST tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。