测试大模型能否像天体物理学家尼尔那样讲科学,发现连最强模型也常答错。
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
- 设计新数据集SCiPS-QA,用742个复杂科学问题测模型回答能力
- 多数开源模型表现远逊于GPT-4 Turbo,但Llama-3-70B表现接近甚至超越
- 连GPT系列也无法可靠验证自身答案,人类也易被错误回复误导
大型语言模型(LLMs)及其驱动的AI助手在专家与普通用户中正快速普及。本文聚焦评估当前大模型作为科学传播者的可靠性。不同于现有基准,本研究关注需深刻理解与答案可回答性意识的科学问答任务。我们提出新数据集SCiPS-QA,包含742个嵌入复杂科学概念的“是/否”问题,并构建评估套件,从正确性与一致性等多维度评测模型。实验涵盖三个来自OpenAI GPT家族的专有模型及13个来自Meta Llama-2、Llama-3和Mistral家族的开源模型。结果显示,多数开源模型显著落后于GPT-4 Turbo,但Llama-3-70B表现突出,多项指标超越GPT-4 Turbo。此外,我们发现即使GPT模型也普遍缺乏可靠验证自身输出的能力。更令人担忧的是,人类评估者常被GPT-4 Turbo的错误回答所误导。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and AI assistants driven by these models are experiencing exponential growth in usage among both expert and amateur users. In this work, we focus on evaluating the reliability of current LLMs as science communicators. Unlike existing benchmarks, our approach emphasizes assessing these models on scientific questionanswering tasks that require a nuanced understanding and awareness of answerability. We introduce a novel dataset, SCiPS-QA, comprising 742 Yes/No queries embedded in complex scientific concepts, along with a benchmarking suite that evaluates LLMs for correctness and consistency across various criteria. We benchmark three proprietary LLMs from the OpenAI GPT family and 13 open-access LLMs from the Meta Llama-2, Llama-3, and Mistral families. While most open-access models significantly underperform compared to GPT-4 Turbo, our experiments identify Llama-3-70B as a strong competitor, often surpassing GPT-4 Turbo in various evaluation aspects. We also find that even the GPT models exhibit a general incompetence in reliably verifying LLM responses. Moreover, we observe an alarming trend where human evaluators are deceived by incorrect responses from GPT-4 Turbo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。