大学生用真实场景测试聊天机器人推理能力,发现其答对但解释常错。
Can Consumer Chatbots Reason? A Student-Led Field Experiment Embedded in an "AI-for-All" Undergraduate Course
- 学生自设计80个推理任务,在真实使用场景中测试主流聊天机器人。
- GPT-5和Claude 4.5答对率最高(86.2%、83.8%),但解释合理性普遍不足。
- 适合教育者用于提升学生AI批判思维,构建可复用的推理测试集。
关于大语言模型聊天机器人是否具备‘推理’能力的讨论,通常依赖精心设计的基准和实验室评估。本文提供一种互补视角:作为乔治梅森大学通识课程UNIV 182(AI-for-All)中期项目,由学生主导的实地实验。学生团队自主设计推理任务,测试当前主流消费者级聊天机器人,并评估答案正确性与推理过程有效性。八支队伍共生成80个原始推理提示,涵盖六类:模式补全、变换规则、空间/视觉推理、定量推理、关系/逻辑推理及类比推理,产生320次模型响应及后续解释。汇总结果显示,OpenAI GPT-5与Claude 4.5平均答对率最高(86.2%、83.8%),其次为Grok 4(82.5%)和Perplexity(73.1%);解释有效性排序类似(81.2%、80.0%、77.5%、66.2%)。定性分析显示,模型在短而结构化的数学与模式题上表现良好,但在空间/视觉推理和多步变换任务中可靠性下降,常见‘答案对但理由错’现象。该实验核心贡献在于教学层面:将人工智能素养转化为实践训练(提示设计、测量、评分分歧、可解释性与真实性),并产出一个基于真实用户交互的学生生成推理探测集。
原文摘要 · Abstract (English)
Claims about whether large language model (LLM) chatbots "reason" are typically debated using curated benchmarks and laboratory-style evaluation protocols. This paper offers a complementary perspective: a student-led field experiment embedded as a midterm project in UNIV 182 (AI4All) at George Mason University, a Mason Core course designed for undergraduates across disciplines with no expected prior STEM exposure. Student teams designed their own reasoning tasks, ran them on widely used consumer chatbots representative of current capabilities, and evaluated both (i) answer correctness and (ii) the validity of the chatbot's stated reasoning (for example, cases where an answer is correct but the explanation is not, or vice versa). Across eight teams that reported standardized scores, students contributed 80 original reasoning prompts spanning six categories: pattern completion, transformation rules, spatial/visual reasoning, quantitative reasoning, relational/logic reasoning, and analogical reasoning. These prompts yielded 320 model responses plus follow-up explanations. Aggregating team-level results, OpenAI GPT-5 and Claude 4.5 achieved the highest mean answer accuracy (86.2% and 83.8%), followed by Grok 4 (82.5%) and Perplexity (73.1%); explanation validity showed a similar ordering (81.2%, 80.0%, 77.5%, 66.2%). Qualitatively, teams converged on a consistent error signature: strong performance on short, structured math and pattern items but reduced reliability on spatial/visual reasoning and multi-step transformations, with frequent "sound right but reason wrong" explanations. The assignment's primary contribution is pedagogical: it operationalizes AI literacy as experimental practice (prompt design, measurement, rater disagreement, and interpretability/grounding) while producing a reusable, student-generated corpus of reasoning probes grounded in authentic end-user interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。