LLMs可生成逻辑题解题提示,但需改进以确保准确性和教学合理性。
The promise and limits of LLMs in constructing proofs and hints for logic problems in intelligent tutoring systems
- 用六种提示法测试四大LLM在358道逻辑题中的逐步推理能力
- DeepSeek-V3达86.7%准确率,简单规则下表现更优
- 生成的提示75%准确且清晰,但解释理由和上下文较弱
智能辅导系统在教授命题逻辑证明方面已证明有效,但依赖模板化解释限制了个性化反馈能力。大型语言模型(LLMs)虽具备动态生成反馈的潜力,却可能产生幻觉或教学不当内容。我们评估了六种提示技术下四种先进LLM在358个命题逻辑问题上的逐步推理准确性。结果显示,DeepSeek-V3在逐步证明构建中达到最高86.7%准确率,尤其在简单规则上表现突出。进一步使用表现最佳的LLM为1,050个学生解题状态生成解释性提示,并通过LLM评分器与人工专家对20%样本进行四维度评估。分析表明,生成提示的准确率为75%,在一致性和清晰度上获人类评委高度评价,但在解释提示原因及整体上下文方面表现不足。结果表明,LLMs可用于增强逻辑辅导系统的提示生成,但需额外改进以保障准确性和教学适切性。
原文摘要 · Abstract (English)
Intelligent tutoring systems have demonstrated effectiveness in teaching formal propositional logic proofs, but their reliance on template-based explanations limits their ability to provide personalized student feedback. While large language models (LLMs) offer promising capabilities for dynamic feedback generation, they risk producing hallucinations or pedagogically unsound explanations. We evaluated the stepwise accuracy of LLMs in constructing multi-step symbolic logic proofs, comparing six prompting techniques across four state-of-the-art LLMs on 358 propositional logic problems. Results show that DeepSeek-V3 achieved superior performance up to 86.7% accuracy on stepwise proof construction and excelled particularly in simpler rules. We further used the best-performing LLM to generate explanatory hints for 1,050 unique student problem-solving states from a logic ITS and evaluated them on 4 criteria with both an LLM grader and human expert ratings on a 20% sample. Our analysis finds that LLM-generated hints were 75% accurate and rated highly by human evaluators on consistency and clarity, but did not perform as well explaining why the hint was provided or its larger context. Our results demonstrate that LLMs may be used to augment tutoring systems with logic tutoring hints, but require additional modifications to ensure accuracy and pedagogical appropriateness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。