arXiv:2603.06878cs.AI2026-03被引 3

LLM回答长短影响人判断,中等长度在错误时更助批判性思维。

Not Too Short, Not Too Long: How LLM Response Length Shapes People's Critical Thinking in Error Detection

  • 实验测试不同长度的LLM解释对用户判断的影响。
  • 错误答案时,中等长度解释使用户正确率最高,正确时各长度表现一致。
  • 适合设计辅助决策系统时参考,尤其关注推理透明度与自信表达。

大型语言模型(LLMs)已成为教育和职场中的常见决策辅助工具,引发对其输出如何影响人类批判性思维的关切。现有研究指出AI协助量会影响认知投入,但对输出具体属性(如回答长度)如何影响用户对信息的评估仍知之甚少。本研究通过24名参与者的被试内实验,考察了在15个修改版沃森-格拉泽批判性思维题目中,带有不同长度和正确性的LLM解释如何影响用户判断准确性。混合效应逻辑回归显示,LLM输出正确性对参与者准确率有显著影响:当解释正确时,用户更可能答对。响应长度在此起调节作用:当解释错误时,中等长度的回应比短或长回应带来更高准确率;而当解释正确时,各长度下的准确率均保持高位。结果表明,仅靠响应长度不足以支撑批判性思维,推理呈现方式(如中等长度在特定条件下更具优势)提示了未来基于LLM的决策支持系统的设计方向,强调推理透明性和确定性表达的平衡。

原文摘要 · Abstract (English)

Large language models (LLMs) have become common decision-support tools across educational and professional contexts, raising questions about how their outputs shape human critical thinking. Prior work suggests that the amount of AI assistance can influence cognitive engagement, yet little is known about how specific properties of LLM outputs (e.g., response length) impacts users' critical evaluation of information. In this study, we examine whether the length of LLM responses shapes users' accuracy in evaluating LLM-generated reasoning on critical thinking tasks, particularly in interaction with the correctness of the LLM's reasoning. To begin evaluating this, we conducted a within-subjects experiment with 24 participants who completed 15 modified Watson--Glaser critical thinking items, each accompanied by an LLM-generated explanation that varied in length and correctness. Mixed-effects logistic regression revealed a strong and statistically reliable effect of LLM output correctness on participant accuracy, with participants more likely to answer correctly when the LLM's explanation was correct. Response length appeared to moderated this effect: when the LLM output was incorrect, medium-length explanations were associated with higher participant accuracy than either shorter or longer explanations, whereas accuracy remained high across lengths when the LLM output was correct. Together, these findings suggest that response length alone may be insufficient to support critical thinking, and that how reasoning is presented-including a potential advantage of mid-length explanations under some conditions-points to design opportunities for LLM-based decision-support systems that emphasize transparent reasoning and calibrated expressions of certainty.

LLM批判性思维用户研究交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。