arXiv:2501.17581cs.CLcs.AI2025-01NAACL被引 9

用自校准大模型实现多维度无参考反仇恨言论评估

CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs

  • 设计自校准思维链的提示方法,自动评分反仇恨言论质量
  • 在四个维度上与人工评价相关性超越传统指标
  • 适合研究自动化内容安全与生成质量评估的学者

反仇恨言论已成为应对网络仇恨言论的有效策略,推动了基于语言模型自动生成反言论的研究。然而,该领域仍缺乏标准化的评估协议和与人类判断一致的可靠自动化评估指标。现有自动评估方法主要依赖相似性度量,无法有效捕捉反言论质量的复杂且独立属性,如语境相关性、攻击性或论证连贯性,导致对人力评估的依赖增加。为此,我们提出CSEval,一个包含四个维度(语境相关性、攻击性、论证连贯性、适用性)的新数据集与评估框架,并引入基于自校准思维链的提示方法Auto-CSEval,用于大语言模型评分。实验表明,Auto-CSEval在与人工评价的相关性上优于ROUGE、METEOR和BertScore等传统指标,显著提升自动化反言论评估能力。

原文摘要 · Abstract (English)

Counterspeech has emerged as a popular and effective strategy for combating online hate speech, sparking growing research interest in automating its generation using language models. However, the field still lacks standardised evaluation protocols and reliable automated evaluation metrics that align with human judgement. Current automatic evaluation methods, primarily based on similarity metrics, do not effectively capture the complex and independent attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence. This has led to an increased dependency on labor-intensive human evaluations to assess automated counter-speech generation methods. To address these challenges, we introduce CSEval, a novel dataset and framework for evaluating counterspeech quality across four dimensions: contextual-relevance, aggressiveness, argument-coherence, and suitableness. Furthermore, we propose Auto-Calibrated COT for Counterspeech Evaluation (Auto-CSEval), a prompt-based method with auto-calibrated chain-of-thoughts (CoT) for scoring counterspeech using large language models. Our experiments show that Auto-CSEval outperforms traditional metrics like ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation.

反仇恨言论评估框架大模型评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。