arXiv:2502.09316cs.CL2025-02被引 4

不依赖人工或大模型评分,用词频统计评估大模型文本生成质量。

A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis

  • 基于n-gram统计和规则构建无裁判评估基准
  • 在50组问答数据上验证,与GPT-4o评估高度一致
  • 计算成本低,适合大规模模型能力评测

评估大语言模型(LLMs)的开放式文本生成能力面临挑战,因缺乏明确的真值且人工或大模型评分成本高昂。本文提出一种新基准,仅使用n-gram统计和规则评估LLMs,无需人类判断或大模型作为裁判。基于50个问题与参考答案对,我们引入三个新指标:流畅性、真实性与帮助性。该基准与GPT-4o评估结果具有强相关性,同时显著降低计算资源消耗,证明其在评估大模型开放式生成能力方面的有效性与可扩展性。

原文摘要 · Abstract (English)

Evaluating the open-ended text generation of large language models (LLMs) is challenging because of the lack of a clear ground truth and the high cost of human or LLM-based assessments. We propose a novel benchmark that evaluates LLMs using n-gram statistics and rules, without relying on human judgement or LLM-as-a-judge approaches. Using 50 question and reference answer sets, we introduce three new metrics based on n-grams and rules: Fluency, Truthfulness, and Helpfulness. Our benchmark strongly correlates with GPT-4o-based evaluations while requiring significantly fewer computational resources, demonstrating its effectiveness as a scalable alternative for assessing LLMs' open-ended generation capabilities.

大模型评估无裁判n-gram生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。