研究发现科学表述的泛化语句在不同群体间理解差异大,易引发误解。
Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models
- 对比普通人、科学家和大模型对科学泛化语句的理解差异。
- 普通人比科学家更认为泛化语句普遍适用且可信,大模型更甚。
- 提醒科研人员与AI在传播中需警惕语言泛化带来的误导风险。
科学家在传达研究结果时常使用泛化表述(如“他汀类药物降低心血管事件”),大语言模型(如ChatGPT-5和DeepSeek)也常沿用此风格。然而,这类未量化陈述易引发过度概括,尤其当不同受众理解不一时。一项比较普通人、科学家及两个主流大模型的研究显示,相较于多数科学家,普通人更倾向于认为这些泛化语句具有广泛适用性和可信度,而大模型则评价更高。这种认知错位揭示了科学传播中的重大风险:科学家可能误以为公众理解一致,而大模型在摘要研究时可能系统性夸大结论。研究强调需重视人类与大模型传播中语言选择的严谨性。
原文摘要 · Abstract (English)
Scientists often use generics, that is, unquantified statements about whole categories of people or phenomena, when communicating research findings (e.g., "statins reduce cardiovascular events"). Large language models (LLMs), such as ChatGPT, frequently adopt the same style when summarizing scientific texts. However, generics can prompt overgeneralizations, especially when they are interpreted differently across audiences. In a study comparing laypeople, scientists, and two leading LLMs (ChatGPT-5 and DeepSeek), we found systematic differences in interpretation of generics. Compared to most scientists, laypeople judged scientific generics as more generalizable and credible, while LLMs rated them even higher. These mismatches highlight significant risks for science communication. Scientists may use generics and incorrectly assume laypeople share their interpretation, while LLMs may systematically overgeneralize scientific findings when summarizing research. Our findings underscore the need for greater attention to language choices in both human and LLM-mediated science communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。