大模型总结科研论文时易过度泛化,可能误导公众理解研究结论。
Generalization Bias in Large Language Model Summarization of Scientific Research
- 测试10个主流大模型,发现多数在总结科研文献时扩大结论范围。
- 新模型比旧模型更易过度泛化,其中3个模型超范围率达26%至73%。
- 相比人类总结,大模型产生宽泛结论的概率高近5倍,适合关注可信度的读者。
由大语言模型(LLMs)驱动的人工智能聊天机器人有潜力提升公众科学素养并支持科研工作,因其能快速将复杂的科学信息转化为通俗易懂的内容。然而,在总结科学文本时,这些模型可能忽略限制性细节,导致结论被过度泛化,超出原研究的合理范围。我们测试了包括ChatGPT-4o、ChatGPT-4.5、DeepSeek、LLaMA 3.3 70B和Claude 3.7 Sonnet在内的10个主流大模型,对比了4900条大模型生成的摘要与原始科学文本。即使明确要求准确性,大多数模型仍存在过度泛化现象,其中DeepSeek、ChatGPT-4o和LLaMA 3.3 70B在26%到73%的案例中出现过度概括。在与人类撰写的科学摘要对比中,大模型摘要包含宽泛推论的可能性是人类的4.85倍(95%置信区间[3.06, 7.70])。值得注意的是,较新的模型在泛化准确性上反而表现更差。结果表明,许多广泛使用的大模型存在显著的过度泛化偏差,可能引发大规模的研究误解。文章提出降低温度参数和建立泛化准确性的基准测试等缓解策略。
原文摘要 · Abstract (English)
Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet, comparing 4900 LLM-generated summaries to their original scientific texts. Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26 to 73% of cases. In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations (OR = 4.85, 95% CI [3.06, 7.70]). Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings. We highlight potential mitigation strategies, including lowering LLM temperature settings and benchmarking LLMs for generalization accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。