arXiv:2605.17382cs.AIcs.CL2026-05

用专家设计的评分标准让AI评估更可靠、可解释且能规模化。

QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

  • 基于多维评分标准,将人类判断与AI评估分离并对齐。
  • 在文本和图像生成任务中,比传统方法更贴近真人评价。
  • 适合需要高可信度评估的AI研发与安全测试场景。

生成式人工智能的快速发展暴露出现有评估方法的根本局限,尤其在开放性、创造性及面向人类的任务中。传统自动指标依赖表面统计相似性,难以反映人类对质量的真实感知;纯人工评估虽可靠但成本高、主观性强且难以扩展。近期使用大语言模型作为评估者的方案提升了可扩展性,却常缺乏对人类定义评估原则的显式支撑,导致偏差与不一致。本文提出量化定性判断(QQJ),一种可扩展且以人为本的评估框架,通过将质量定义与执行分离,以专家设计的多维度评分标准为锚点,并利用少量高质量标注数据校准大语言模型评估者,使其推理与专家一致。该设计实现了跨多种生成任务与模态的一致性、可解释性和可扩展性评估。在文本与图像生成上的大量实验表明,QQJ在与人类判断的对齐程度上显著优于传统自动指标和无约束的LLM评估者。此外,其在重复评估中表现出更强稳定性,并具备更强的故障诊断能力,可有效识别幻觉与意图错位等关键问题。结果表明,结构化定性判断可在不牺牲可解释性或人类对齐的前提下实现规模化,为现代生成式AI系统提供可靠的评估基础。

原文摘要 · Abstract (English)

The rapid progress of generative artificial intelligence has exposed fundamental limitations in existing evaluation methodologies, particularly for open-ended, creative, and human-facing tasks. Traditional automatic metrics rely on surface-level statistical similarity and often fail to reflect human perceptions of quality, while purely human evaluation, although reliable, is costly, subjective, and difficult to scale. Recent approaches using large language models as evaluators offer improved scalability but frequently lack explicit grounding in human-defined evaluation principles, leading to bias and inconsistency. In this paper, we introduce Quantifying Qualitative Judgment (QQJ), a scalable and human-centric evaluation framework that explicitly bridges the gap between human judgment and automated assessment. QQJ separates the definition of quality from its execution by anchoring evaluation in expert-designed, multi-dimensional rubrics and calibrating large language model evaluators to align with expert reasoning using a small, high-quality annotation set. This design enables consistent, interpretable, and scalable evaluation across diverse generative tasks and modalities. Extensive experiments on text and image generation demonstrate that QQJ achieves substantially stronger alignment with human judgment than traditional automatic metrics and unconstrained LLM-based evaluators. Moreover, QQJ exhibits improved stability across repeated evaluations and superior diagnostic capability in identifying critical failure modes such as hallucination and intent mismatch. These results indicate that structured qualitative judgment can be operationalized at scale without sacrificing interpretability or human alignment, positioning QQJ as a practical foundation for reliable evaluation of modern generative AI systems.

AI评估人机对齐大模型评测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。