arXiv:2605.06054cs.AIcs.HC2026-05

用视觉指纹对比大模型在不同设置下的生成风格差异。

Visual Fingerprints for LLM Generation Comparison

论文配图:Visual Fingerprints for LLM Generation Comparison
图 1 · 摘自论文原文
  • 将文本生成视为语言选择的集合,建模其分布特征。
  • 通过重复采样生成响应,提取并可视化分布模式。
  • 适合关注提示设计与模型行为分析的研究者使用。

大语言模型输出受提示、系统指令、模型参数和架构等多种因素复杂交互影响。我们将这些因素的具体配置称为生成条件,每种条件都会以不同方式影响输出结果。理解不同生成条件如何塑造模型行为,对提示设计和模型评估至关重要,但因文本生成的随机性和开放性而难以实现。本文提出一种通过建模响应中内容、表达和结构等语言选择的分布,来可视化比较不同生成条件下大模型输出的方法。我们利用自然语言处理流程提取这些选择,并表示为分布形式,进而生成可视化的‘语言指纹’,实现条件特异性倾向的分布级直接对比。在四个应用场景中,我们展示了视觉指纹能够揭示出单个样本或平均指标难以察觉的一致性行为模式。

原文摘要 · Abstract (English)

Large language model (LLM) outputs arise from complex interactions among prompts, system instructions, model parameters, and architecture. We refer to specific configurations of these factors as generation conditions, each of which can bias outputs in various ways. Understanding how different generation conditions shape model behaviors is essential for tasks such as prompt design and model evaluation, yet it remains challenging due to the stochastic and open-ended nature of text generation. We present an approach to visually compare LLM outputs across generation conditions by modeling responses as collections of linguistic choices, including content, expression, and structure. We extract these choices using natural language processing pipelines and represent their distributions across repeated samples. We then visualize these distributions as visual fingerprints, enabling direct, distribution-level comparison of condition-specific tendencies. Through four usage scenarios, we demonstrate how visual fingerprints reveal consistent patterns in LLM behavior that are difficult to observe through individual responses or aggregate metrics.

大模型分析生成对比可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。