arXiv:2512.16323cs.CL2025-12中稿 · EACL2026 main被引 1

用一句话就能骗过神经文本评估模型,暴露其脆弱性。

Hacking Neural Evaluation Metrics with Single Hub Text

  • 在离散空间中寻找单个欺骗性文本,使其始终被高分评价。
  • 该文本在日英、德英翻译任务中分别获79.1%和67.8% COMET分。
  • 结果超越通用模型M2M100的个别生成结果,且跨语言泛化有效。

强相关于人类判断的评估指标是生成模型发展与改进的重要指南,必须具备高度可靠性与鲁棒性。近年来基于嵌入的神经文本评估指标(如用于翻译任务的COMET)在科研与工业界广泛应用。然而,由于神经网络的黑箱特性,这些指标的评估结果缺乏保障。为揭示此类指标的可靠性与安全性问题,我们提出一种方法,在离散空间中寻找一个单一对抗性文本,无论测试样本如何,均被一致评为高质量,以识别评估指标的漏洞。使用该方法找到的“枢纽文本”在WMT'24英语到日语(En--Ja)和英语到德语(En--De)翻译任务中分别取得79.1%和67.8%的COMET得分,优于使用M2M100通用翻译模型对每个源句单独生成的结果。此外,我们还验证了该枢纽文本在多个语言对(如日英、德英)之间具有良好的泛化能力。

原文摘要 · Abstract (English)

Strongly human-correlated evaluation metrics serve as an essential compass for the development and improvement of generation models and must be highly reliable and robust. Recent embedding-based neural text evaluation metrics, such as COMET for translation tasks, are widely used in both research and development fields. However, there is no guarantee that they yield reliable evaluation results due to the black-box nature of neural networks. To raise concerns about the reliability and safety of such metrics, we propose a method for finding a single adversarial text in the discrete space that is consistently evaluated as high-quality, regardless of the test cases, to identify the vulnerabilities in evaluation metrics. The single hub text found with our method achieved 79.1 COMET% and 67.8 COMET% in the WMT'24 English-to-Japanese (En--Ja) and English-to-German (En--De) translation tasks, respectively, outperforming translations generated individually for each source sentence by using M2M100, a general translation model. Furthermore, we also confirmed that the hub text found with our method generalizes across multiple language pairs such as Ja--En and De--En.

评估漏洞对抗攻击COMET翻译评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。