arXiv:2608.17731cs.AI2026-08

用多样性曲线取代单一分数,更真实评估AI生成内容的丰富程度。

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

  • 提出多样性曲线,动态评估不同参数下的生成内容差异
  • 实验证明单一数值指标常因参数选择导致矛盾排名
  • 适合关注生成内容多样性的研究人员和评测者

多样性是评估生成式人工智能系统的核心标准,但其度量仍存在本质模糊性。现有方法通常将生成样本映射到嵌入空间,计算成对距离或相似性,并聚合为单一标量分数。这类标量总结虽便捷,却蕴含不同归纳偏置,常导致同一样本集产生矛盾排序。本文认为,将多样性评估简化为单一数值本质上是不充分的。我们首先回顾代表性多样性度量,从两个互补角度诊断其局限:公理化分析表明,不存在满足所有理想性质的标量度量;实证分析揭示高维表示空间会引发集中且依赖模态的距离分布。为此,我们提出多样性曲线——一种条件感知的曲线型摘要,可在指定表示和距离/核函数下,评估参数化多样性族在多个阈值、尺度、指数或阶数下的表现。多样性曲线揭示比较结果是否对分辨率鲁棒,还是依赖于任意参数选择。我们实例化了多种代表性度量家族的曲线,并展示了其在生成式AI评估中的实际应用。总体而言,多样性曲线提供了一个更透明、分辨率感知的框架,用于比较AI生成内容的多样性。

原文摘要 · Abstract (English)

Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that diversity evaluation for AI-generated content is intrinsically under-specified when reduced to a single number. We first review representative diversity metrics, and then diagnose their limitations from two complementary perspectives: an axiomatic analysis showing that no representative scalar metric satisfies all desirable properties simultaneously, and an empirical analysis showing that high-dimensional representation spaces can induce concentrated, modality-dependent distance distributions. To address these issues, we propose diversity profiles: curve-valued, condition-aware summaries that evaluate a parameterized diversity family across a range of thresholds, scales, exponents, or orders under a specified representation and distance or kernel function. Diversity profiles reveal whether a comparison is robust across resolutions or instead depends on an arbitrary parameter choice. We instantiate profiles for several representative metric families and demonstrate their practical use in generative AI evaluation. Overall, diversity profiles provide a more transparent and resolution-aware framework for comparing the diversity of AI-generated content.

多样性评估生成模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。