评测大模型有害性评估指标,发现传统方法竟比大模型自身更准。
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- 构建统一基准HarmMetric Eval,涵盖多种有害内容类别。
- 实验证明ROUGE、METEOR等传统指标在细粒度评估中优于大模型裁判。
- 提出新裁判设计:结合细粒度标准与轻量微调,效果领先。
大型语言模型(LLMs)生成有害内容的潜在风险对数据管理构成重大威胁,因其被广泛用作数据生成引擎。尽管已有众多有害性评估指标和裁判被提出,但因格式与量表差异,它们在评估模型生成的有害内容时结果不一致,削弱了实际可信度。为此,本文提出HarmMetric Eval,一个系统性基准,用于评估不同格式与量表的有害性指标与裁判质量。该基准包含高质量数据集,涵盖多细粒度类别的代表性有害提示及其对应的有害与非有害输出,并采用统一评分机制,奖励能正确排序有害与非有害输出的指标。大量实验揭示惊人发现:传统基于参考的指标如ROUGE和METEOR,在细粒度有害性评估中表现优于基于LLM的裁判,挑战了主流认为大模型更优的假设。通过细粒度分析,我们解释了大模型裁判在评估无关或无用输出时的局限性。受此启发,我们设计了一种改进型有害性裁判,其提示模板明确融入细粒度有害性标准,并利用参考指标对基础大模型进行轻量级微调。该裁判在HarmMetric Eval上达到当前最优评估效果。
原文摘要 · Abstract (English)
The potential of large language models (LLMs) to generate harmful content poses a significant safety risk for data management, as LLMs are increasingly being used as engines for data generation. To assess this risk, numerous harmfulness evaluation metrics and judges have been proposed. However, due to differences in their formats and scales, these metrics may yield inconsistent evaluation results on LLM-generated harmful data, undermining their credibility in practice. To address this gap, we present HarmMetric Eval, a systematic benchmark for assessing the quality of harmfulness metrics and judges with varying formats and scales. HarmMetric Eval includes a high-quality dataset comprising representative harmful prompts paired with harmful and non-harmful LLM outputs across multiple fine-grained categories, along with a unified scoring mechanism to reward the metrics for correctly ranking harmful outputs over non-harmful ones. Extensive experiments on HarmMetric Eval yield a surprising finding: conventional reference-based metrics such as ROUGE and METEOR can outperform LLM-based judges in fine-grained harmfulness evaluation, challenging prevailing assumptions about LLMs' superiority in this domain. To reveal the reasons behind this finding, we provide a fine-grained analysis to explain the limitations of LLM-based judges on rating irrelevant or useless LLM outputs. Motivated by these insights, we design an improved harmfulness judge that explicitly incorporates fine-grained harmfulness criteria in its prompt template and leverages reference-based metrics for lightweight fine-tuning of its base LLM. The resulting judge achieves state-of-the-art evaluation effectiveness on HarmMetric Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。