arXiv:2604.18724cs.AI2026-04

让大模型生成结果的分布可见,帮助用户看清输出多样性与敏感性。

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations

  • 用路径图可视化多个生成结果,展现共性结构与分支点。
  • 实验表明图形总结更利于判断多样性,逐条查看更适合细节分析。
  • 适合需要理解生成模式的研究者和提示工程师使用。

用户通常通过单一输出来交互和评估语言模型,但每个输出只是众多可能完成中的一个样本。这种交互方式隐藏了分布结构,如模式、罕见边缘情况以及对微小提示变化的敏感性,导致用户在开放式任务中迭代提示时容易过度概括。基于对13位使用大模型的研究人员的形成性研究,探讨了随机性在实践中何时重要、如何推理语言分布以及现有工作流的断裂点,我们提出GROVE。GROVE是一种交互式可视化工具,将多个语言模型生成结果表示为文本图中的重叠路径,揭示共享结构、分支点和聚类,同时保留原始输出的可访问性。我们在三个众包用户研究中评估(参与人数分别为47、44和40),针对互补的分布任务。结果支持一种混合工作流:图示摘要有助于结构判断(如评估多样性),而直接查看输出在细节问题上仍更有效。

原文摘要 · Abstract (English)

Users typically interact with and evaluate language models via single outputs, but each output is just one sample from a broad distribution of possible completions. This interaction hides distributional structure such as modes, uncommon edge cases, and sensitivity to small prompt changes, leading users to over-generalize from anecdotes when iterating on prompts for open-ended tasks. Informed by a formative study with researchers who use LMs (n=13) examining when stochasticity matters in practice, how they reason about distributions over language, and where current workflows break down, we introduce GROVE. GROVE is an interactive visualization that represents multiple LM generations as overlapping paths through a text graph, revealing shared structure, branching points, and clusters while preserving access to raw outputs. We evaluate across three crowdsourced user studies (N=47, 44, and 40 participants) targeting complementary distributional tasks. Our results support a hybrid workflow: graph summaries improve structural judgments such as assessing diversity, while direct output inspection remains stronger for detail-oriented questions.

语言模型可视化生成分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。