实验证明,大模型的思维链推理在智能体流水线中无法提升效果或解释力。
Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
- 用多模型协作的智能体流水线测试思维链推理机制
- 思维链未提升输出质量,也未增强用户理解能力
- 适合关注AI可解释性局限的研究者和实践者
智能体流水线为以人为中心的可解释性带来新挑战与机遇。当前HCXAI社区仍难以以实用方式揭示大模型内部运作。本研究展示了一个感知任务引导系统的智能体流水线实现。通过定量与定性分析,考察了思维链(Chain-of-Thought, CoT)推理在该场景中的表现。结果表明,仅依赖思维链推理并不能提升输出质量,也无法提供真正有效的可解释性——其生成的解释往往缺乏实际意义,未能增强终端用户对系统行为的理解或帮助其实现目标。
原文摘要 · Abstract (English)
Agentic pipelines present novel challenges and opportunities for human-centered explainability. The HCXAI community is still grappling with how best to make the inner workings of LLMs transparent in actionable ways. Agentic pipelines consist of multiple LLMs working in cooperation with minimal human control. In this research paper, we present early findings from an agentic pipeline implementation of a perceptive task guidance system. Through quantitative and qualitative analysis, we analyze how Chain-of-Thought (CoT) reasoning, a common vehicle for explainability in LLMs, operates within agentic pipelines. We demonstrate that CoT reasoning alone does not lead to better outputs, nor does it offer explainability, as it tends to produce explanations without explainability, in that they do not improve the ability of end users to better understand systems or achieve their goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。