评测大模型识别荒谬科学问题的能力,提出生成合成数据的新方法。
SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation
- 用受控的生成方法构造含逻辑谬误的科学问题数据集
- 发现80%测试中模型仍会答出0.5个孩子等荒谬答案
- 揭示模型认知波动性,适合评估AI可信度的研究者使用
当前大语言模型(如GPT-4o、GPT-o1-preview、Gemini Flash)在面对‘一人一女一年生一孩,三人一女半年生几孩’这类荒谬问题时,常错误回答‘0.5’。尽管部分模型能识别问题不实,但在10次试验中有8次仍给出不合逻辑的答案。我们观察到:若模型曾正确识别问题缺陷,后续响应更可能保持一致理解,但该能力不稳定。为此,我们构建了名为SciFaultyQA的科学问题数据集,其中问题故意设计为逻辑或科学上不合理。通过分析模型在这些题目上的表现,提出一种受生成对抗网络启发的合成数据生成方法,用于系统评估和基准测试不同大模型在识别此类故障问题上的能力,并开发新策略降低错误率。
原文摘要 · Abstract (English)
Consider the problem: ``If one man and one woman can produce one child in one year, how many children will be produced by one woman and three men in 0.5 years?" Current large language models (LLMs) such as GPT-4o, GPT-o1-preview, and Gemini Flash frequently answer "0.5," which does not make sense. While these models sometimes acknowledge the unrealistic nature of the question, in many cases (8 out of 10 trials), they provide the nonsensical answer of "0.5 child." Additionally, temporal variation has been observed: if an LLM answers correctly once (by recognizing the faulty nature of the question), subsequent responses are more likely to also reflect this understanding. However, this is inconsistent. These types of questions have motivated us to develop a dataset of science questions, SciFaultyQA, where the questions themselves are intentionally faulty. We observed that LLMs often proceed to answer these flawed questions without recognizing their inherent issues, producing results that are logically or scientifically invalid. By analyzing such patterns, we developed a novel method for generating synthetic datasets to evaluate and benchmark the performance of various LLMs in identifying these flawed questions. We have also developed novel approaches to reduce the errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。