arXiv:2504.00938cs.AIcs.LG2025-04被引 2

用统计方法验证AI能否像专家一样评设计,结果表明带推理的AI表现不输人类专家。

AI Judges in Design: Statistical Perspectives on Achieving Human Expert Equivalence With Vision-Language Models

  • 构建统计框架,量化AI评分与人类专家的一致性。
  • 最优AI模型在4项设计指标中3项达到专家水平,6次测试中有5次超越多数新手。
  • 适合教育、工业设计等领域快速评估创意作品,也适用于其他主观评价场景。

早期工程设计(如概念草图)的主观评估传统依赖人类专家,但耗时、昂贵且易不一致。视觉-语言模型(VLMs)为自动化评估带来可能,但需确保其表现与人类专家相当。本文提出一个严谨的统计框架,用于判断AI“评委”是否达到人类专家等效。案例研究评估了四种基于VLM的裁判,在独特性、创造性、实用性及绘图质量四项关键指标上的表现。这些AI裁判采用不同上下文学习(ICL)技术,包括单模态与多模态提示及推理能力。同一框架也用于评估三位训练过的新人是否具备专家等效性。结果显示,使用文本与图像结合的ICL并含推理的最优AI裁判,在独特性和绘图质量上达成专家级一致性;在全部四项指标上表现优于或匹配训练新人。在6次运行中,独特性与创造性均达6次以上,绘图质量与实用性各达5次以上,其与专家的一致性超过多数训练新人。结果表明,支持推理的VLM模型可实现设计评估的人类专家等效。该方法对教育与实践中规模化设计评估具有意义,并提供通用统计框架以验证其他领域中的主观内容评估AI裁判。

原文摘要 · Abstract (English)

The subjective evaluation of early stage engineering designs, such as conceptual sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these AI ``judges'' perform on par with human experts. However, no existing framework assesses expert equivalence. This paper introduces a rigorous statistical framework to determine whether an AI judge's ratings match those of human experts. We apply this framework in a case study evaluating four VLM-based judges on key design metrics (uniqueness, creativity, usefulness, and drawing quality). These AI judges employ various in-context learning (ICL) techniques, including uni- vs. multimodal prompts and inference-time reasoning. The same statistical framework is used to assess three trained novices for expert-equivalence. Results show that the top-performing AI judge, using text- and image-based ICL with reasoning, achieves expert-level agreement for uniqueness and drawing quality and outperforms or matches trained novices across all metrics. In 6/6 runs for both uniqueness and creativity, and 5/6 runs for both drawing quality and usefulness, its agreement with experts meets or exceeds that of the majority of trained novices. These findings suggest that reasoning-supported VLM models can achieve human-expert equivalence in design evaluation. This has implications for scaling design evaluation in education and practice, and provides a general statistical framework for validating AI judges in other domains requiring subjective content evaluation.

AI评审设计评估视觉语言模型统计验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。