arXiv:2509.01412cs.CL2025-09

让人类实时干预大模型推理过程,提升准确率与可信度

Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning

  • 将线性推理文本转为可交互逻辑图,支持可视化追踪
  • 在GSM8K和StrategyQA上准确率最高提升24个百分点
  • 适合需要高可靠性的决策场景,如医疗、金融

大型语言模型通过思维链(CoT)提示展现强大推理能力,但过程不透明,难以在高风险场景中进行验证、调试和控制。我们提出Vis-CoT,一个基于人机协同的框架,将线性CoT文本转化为可交互的推理图。用户可可视化逻辑流程,识别错误步骤,并通过剪枝错误路径、嫁接用户自定义前提进行干预。这一机制将交互从被动观察转变为主动协作,引导模型得出更准确、可信的结论。在GSM8K和StrategyQA数据集上,Vis-CoT相比非交互基线,最终答案准确率提升最高达24个百分点。用户研究进一步显示,其在可用性和信任感方面均有显著提升。Vis-CoT为结合大模型与针对性人类监督,实现更可靠、可理解、协作式推理提供了可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) show strong reasoning via chain-of-thought (CoT) prompting, but the process is opaque, which makes verification, debugging, and control difficult in high-stakes settings. We present Vis-CoT, a human-in-the-loop framework that converts linear CoT text into an interactive reasoning graph. Users can visualize the logical flow, identify flawed steps, and intervene by pruning incorrect paths and grafting new, user-defined premises. This shifts interaction from passive observation to active collaboration, steering models toward more accurate and trustworthy conclusions. Across GSM8K and StrategyQA, Vis-CoT improves final-answer accuracy by up to 24 percentage points over non-interactive baselines. A user study also shows large gains in perceived usability and trust. Vis-CoT points to a practical path for more reliable, understandable, and collaborative reasoning by combining LLMs with targeted human oversight.

人机协同思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。