arXiv:2606.05983cs.AIcs.CL2026-06被引 1

提出FJS三维度模型,评估学生用生成式AI思考的能力

Framing, Judging, Steering: An Assessable Competency Model for Teach-ing Students to Reason With Generative AI

  • 将人机协作思维拆解为设问、判断、调整三步,独立可测
  • 模拟学习者实验显示三能力可分离,评分具一致性和区分度
  • 适合教育研究者与教师,用于设计和评估AI素养教学

生成式AI让答案触手可及,却使理解变得困难,盲目使用会引发认知卸载。学校仍考核无辅助表现,但真实任务是借助AI产出优质成果:明确定义模糊问题、评估输出错误与隐含假设、迭代引导模型改进。这一能力极少被单独评估;现有测评常合并为单一“提示”得分,无法诊断成败原因。本文提出CoRe-3(协同推理)能力模型,将有效使用AI分解为三个可评估技能:FJS——Framing(设问,即调用AI前明确模糊任务)、Judging(判断,即评估输出中的错误与未言明假设)、Steering(调整,即迭代重定向模型)。其核心主张是将生成前的设问与生成后的调整分离,以判断作为中间关口。模型基于理论构建,提出五个可检验命题,并在开放平台CoReasoningLab中实现,该平台呈现有缺陷的AI输出并独立评分。对模拟学习者的测试表明,三能力可分离:每项仅随自身操控而变化,其余保持不变;当三项共享同一能力时,评分相关性上升(收敛效度与区分效度),且跨两家评分模型提供商均成立。下一步将开展真人评分一致性与实际效果验证。本文发布工具、数据与流程。

原文摘要 · Abstract (English)

Generative AI makes answers easy and understanding hard, and uncritical use invites cognitive offloading. Schools still measure unaided performance, yet the real task is to produce good work with AI: framing an ill-defined task, judging the output, and steering the model toward a better result. This ability is rarely assessed in its own right; where measured, it collapses into one "prompting" score that cannot diagnose why AI use succeeds or fails. We propose CoRe-3 (Co-Reasoning), a competency model factoring productive AI use into three assessable skills we abbreviate FJS: Framing (specifying an ill-defined task before invoking AI), Judging (evaluating output for errors and unstated assumptions), and Steering (iteratively redirecting the model). Its distinguishing claim is the separation of pre-generation Framing from post-generation Steering, with Judging as the gate between. We ground the skills in theory, state five testable propositions, and instantiate them in CoReasoningLab, an open platform that presents flawed AI output and scores them independently. Over simulated learners (generated and graded by different models), the skills dissociate: each tracks its own manipulated competence while staying flat in the others, and grades become correlated when one competence is shared across all three (convergent and discriminant validity), across grader backends from two providers. Human-rater agreement and outcomes are next; we release the instrument, data, and protocol.

AI教育评估模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。