arXiv:2503.17482cs.LGcs.AI2025-03NeurIPS被引 4

评估生成模型时,能生成好内容不等于用户能精准控制产出。

What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models

  • 提出可独立于生成能力衡量可控性的新指标:可引导性。
  • 用户实验显示,主流模型在精准复现目标输出上表现差(<20%成功)。
  • 图像引导等简单方法能让可引导性提升2倍以上,适合交互设计研究者。

如何评估生成模型的质量?现有指标多关注模型的可生成性,即其输出的质量与多样性。但生成模型的实际价值不仅在于它能产生什么,更在于用户能否根据特定目标生成满足需求的输出——我们称此为可引导性。本文首次提出数学分解方法,独立量化可引导性。由于需知晓用户目标,可引导性评估更具挑战性,为此我们设计了基于用户复现生成样本的基准任务。在文本到图像及大语言模型的用户研究中发现,尽管这些模型生成质量高,但可引导性普遍偏低(平均成功率低于20%)。结果表明应重视可引导性改进。我们验证了其可行性:简单的图像引导机制使该基准性能提升超过2倍。

原文摘要 · Abstract (English)

How should we evaluate the quality of generative models? Many existing metrics focus on a model's producibility, i.e. the quality and breadth of outputs it can generate. However, the actual value from using a generative model stems not just from what it can produce but whether a user with a specific goal can produce an output that satisfies that goal. We refer to this property as steerability. In this paper, we first introduce a mathematical decomposition for quantifying steerability independently from producibility. Steerability is more challenging to evaluate than producibility because it requires knowing a user's goals. We address this issue by creating a benchmark task that relies on one key idea: sample an output from a generative model and ask users to reproduce it. We implement this benchmark in user studies of text-to-image and large language models. Despite the ability of these models to produce high-quality outputs, they all perform poorly on steerability. These results suggest that we need to focus on improving the steerability of generative models. We show such improvements are indeed possible: simple image-based steering mechanisms achieve more than 2x improvement on this benchmark.

生成模型可引导性用户研究评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。