VISTA让大模型多轮推理过程可视觉化,轻松看清逻辑链条。
Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
- 通过交互式可视化分析对话上下文对模型决策的影响
- 自动生成推理依赖树,清晰展示每步逻辑路径
- 支持多种模型和评测集,适合研究者调试与对比
近期研究日益关注大语言模型在多轮交互中的推理能力,这类场景更贴近真实问题解决。然而,由于复杂的上下文依赖和缺乏专用可视化工具,分析这些交互中的推理过程面临巨大挑战,给研究人员带来高认知负担。为此,我们提出 VISTA——一个基于网页的多轮推理任务文本分析可视化交互系统。VISTA 允许用户可视化上下文对模型决策的影响,并可交互修改对话历史,进行跨模型的「假设分析」。此外,平台能自动解析会话并生成推理依赖树,提供模型逐步推理过程的透明视图。通过统一且交互式的框架,VISTA 显著降低推理链分析的复杂度,促进对当前 LLM 能力与局限的深入理解。该平台开源,支持自定义基准测试和本地模型集成。
原文摘要 · Abstract (English)
Recent research has increasingly focused on the reasoning capabilities of Large Language Models (LLMs) in multi-turn interactions, as these scenarios more closely mirror real-world problem-solving. However, analyzing the intricate reasoning processes within these interactions presents a significant challenge due to complex contextual dependencies and a lack of specialized visualization tools, leading to a high cognitive load for researchers. To address this gap, we present VISTA, an web-based Visual Interactive System for Textual Analytics in multi-turn reasoning tasks. VISTA allows users to visualize the influence of context on model decisions and interactively modify conversation histories to conduct "what-if" analyses across different models. Furthermore, the platform can automatically parse a session and generate a reasoning dependency tree, offering a transparent view of the model's step-by-step logical path. By providing a unified and interactive framework, VISTA significantly reduces the complexity of analyzing reasoning chains, thereby facilitating a deeper understanding of the capabilities and limitations of current LLMs. The platform is open-source and supports easy integration of custom benchmarks and local models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。