arXiv:2502.04718cs.CL2025-02NAACL被引 6

评估文本风格迁移的可靠度量,发现大模型和新指标更有效。

Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?

  • 对比多种自动评价指标,测试其在情感迁移与去毒化任务中的表现
  • 在英、印地语、孟加拉语三语场景下,大模型评价比传统指标相关性更高
  • 集成多个指标或使用大模型可显著提升评价可靠性,适合研究者参考

文本风格迁移(TST)旨在改变文本风格的同时保留原始内容。其评估涉及风格转移准确性、内容保留性和自然度等多个维度。人工评价虽理想但成本高,而针对TST的自动评价指标研究远不如机器翻译或摘要等任务充分。本文在英语、印地语和孟加拉语背景下,考察了来自更广泛NLP任务的现有及新型指标在两个典型子任务——情感迁移与去毒化中的表现。通过与人类判断进行元评估(meta-evaluation),我们验证了这些指标单独或组合使用时的有效性。此外,我们探索了大语言模型(LLMs)作为评价工具的潜力。结果表明,新应用的先进NLP指标及基于大模型的评估方法优于现有TST评价指标,而最优的oracle集成方法展现出更大潜力。

原文摘要 · Abstract (English)

Text style transfer (TST) is the task of transforming a text to reflect a particular style while preserving its original content. Evaluating TST outputs is a multidimensional challenge, requiring the assessment of style transfer accuracy, content preservation, and naturalness. Using human evaluation is ideal but costly, as is common in other natural language processing (NLP) tasks, however, automatic metrics for TST have not received as much attention as metrics for, e.g., machine translation or summarization. In this paper, we examine both set of existing and novel metrics from broader NLP tasks for TST evaluation, focusing on two popular subtasks, sentiment transfer and detoxification, in a multilingual context comprising English, Hindi, and Bengali. By conducting meta-evaluation through correlation with human judgments, we demonstrate the effectiveness of these metrics when used individually and in ensembles. Additionally, we investigate the potential of large language models (LLMs) as tools for TST evaluation. Our findings highlight newly applied advanced NLP metrics and LLM-based evaluations provide better insights than existing TST metrics. Our oracle ensemble approaches show even more potential.

风格迁移评估指标大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。