让大模型生成文本像文献一样可细读,支持多维度对比分析。
LLMbench: A Comparative Close Reading Workbench for Large Language Models
- 双模型输出并列显示,支持词级差异、语气、结构等四类分析叠加
- 提供六种分析模式,揭示生成文本背后的概率分布与可能路径
- 适合人文社科研究者批判性分析大模型生成内容的工具
LLMbench 是一个基于浏览器的大语言模型输出比较细读工作台。不同于侧重量化评估与用户评分的现有工具(如 Google PAIR 的 LLM Comparator),LLMbench 聚焦数字人文领域的诠释学实践。同一提示下两个模型的输出并列展示在可注释面板中,配备四种分析叠加层:逐标记概率(用于检查词级对数概率)、词级差异(对比两文本)、语气(基于 Hyland 风格的元话语分析)和结构(句级解析,突出连贯词)。同时提供五种分析模式:随机变异、温度梯度、提示敏感性、标记概率和跨模型分歧,使生成文本在标记层面的概率结构可视化。该工具将生成文本视为独立的研究对象,从概率分布视角呈现其潜在的反事实历史,提供连续热图、熵值火花线、像素图和三维概率地形等可视化手段,展现每个词的生成轨迹。本文阐述了工具架构、六种模式及其设计动机,并主张当前在人文学科和社会科学中被忽视的对数概率数据,是开展生成式 AI 批判研究的重要资源。
原文摘要 · Abstract (English)
LLMbench is a browser-based workbench for the comparative close reading of large language model (LLM) outputs. Where existing tools for LLM comparison, such as Google PAIR's LLM Comparator are engineered for quantitative evaluation and user-rating metrics, LLMbench is oriented towards the hermeneutic practices of the digital humanities. Two model responses to the same prompt are side by side in annotatable panels with four analytical overlays (Probabilities for token-level log-probability inspection, Differences for word-level diff across the two panels, Tone for Hyland-style metadiscourse analysis, and Structure for sentence-level parsing with discourse connective highlighting), alongside five analytical modes, Stochastic Variation, Temperature Gradient, Prompt Sensitivity, Token Probabilities, and Cross-Model Divergence, that make the probabilistic structure of generated text legible at the token level. The tool treats the generated text as a research object in its own right from a probability distribution, a text that could have been otherwise, and provides visualisations including continuous heatmaps, entropy sparklines, pixel maps, and three-dimensional probability terrains, that show the counterfactual history from which each word emerged. This paper describes the tool's architecture, its six modes, and its design rationale, and argues that log-probability data, currently underused in humanistic and social-scientific readings of AI, is an important resource for a critical studies of generative AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。