arXiv:2501.06728cs.CL2025-01被引 1

测试无参考对话评估系统在四种攻击下的鲁棒性

Measuring the Robustness of Reference-Free Dialogue Evaluation Systems

  • 构建对抗性攻击基准,测试无参考评估指标的稳定性
  • 发现传统相关性高的指标在攻击下表现差异显著
  • 适合关注对话评估可靠性的研究者和开发者

基于大语言模型的对话系统进展迅速,但评估指标发展滞后,尤其在多样性和创造性回复方面。本文提出一个基准,用于评估无参考对话评估指标对四类对抗性攻击的鲁棒性:说话人标签前缀、静态回复、语法错误回复和重复对话上下文。我们分析了DialogRPT、UniEval和PromptEval(一种基于提示的LLM方法)在有事实依据和无事实依据数据集上的表现。通过考察指标与人类判断的相关性及对对抗攻击的敏感度,发现这两者并不总是一致;传统基准中表现相近的指标,在对抗性回复下评分差异明显。这一结果推动了更精细评估框架的发展,以应对真实对话场景中的挑战。

原文摘要 · Abstract (English)

Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluating the robustness of reference-free dialogue metrics against four categories of adversarial attacks: speaker tag prefixes, static responses, ungrammatical responses, and repeated conversational context. We analyze metrics such as DialogRPT, UniEval, and PromptEval -- a prompt-based method leveraging LLMs -- across grounded and ungrounded datasets. By examining both their correlation with human judgment and susceptibility to adversarial attacks, we find that these two axes are not always aligned; metrics that appear to be equivalent when judged by traditional benchmarks may, in fact, vary in their scores of adversarial responses. These findings motivate the development of nuanced evaluation frameworks to address real-world dialogue challenges.

对话评估鲁棒性大模型无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。