解决大模型在数学题纠错中的从众偏差问题,提升检测准确率。
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
- 先生成多种解题参考路径,再进行错误检测,避免单一答案干扰。
- 在GSM8K数据集上测试,结合思维链提示后性能显著提升。
- 适合教育AI、自动批改系统开发者使用,尤其关注多解题场景。
大语言模型为数学应用题的自动错误检测带来新机遇。现有研究虽证实其有效性,却忽略了单个题目可能存在多个正确解法的问题。我们初步分析发现,常规解法与替代解法之间存在显著性能差异,这一现象称为“从众偏差”。为此,我们提出Ask-Before-Detect(AskBD)框架,利用大模型生成自适应的参考解法,以增强错误检测能力。在200个GSM8K样例上的实验表明,AskBD能有效缓解偏差,尤其当结合思维链提示等推理增强技术时效果更优。
原文摘要 · Abstract (English)
The rise of large language models (LLMs) offers new opportunities for automatic error detection in education, particularly for math word problems (MWPs). While prior studies demonstrate the promise of LLMs as error detectors, they overlook the presence of multiple valid solutions for a single MWP. Our preliminary analysis reveals a significant performance gap between conventional and alternative solutions in MWPs, a phenomenon we term conformity bias in this work. To mitigate this bias, we introduce the Ask-Before-Detect (AskBD) framework, which generates adaptive reference solutions using LLMs to enhance error detection. Experiments on 200 examples of GSM8K show that AskBD effectively mitigates bias and improves performance, especially when combined with reasoning-enhancing techniques like chain-of-thought prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。