arXiv:2412.05200cs.AI2024-12

测试大模型在科学馆问答表现,发现创意与准确难兼顾。

Are Frontier Large Language Models Suitable for Q&A in Science Centres?

  • 用三个主流大模型生成儿童友好的问答响应
  • 克劳德在准确性和吸引力上优于GPT和Gemini
  • 越有创意的回答,事实可靠性越低,需谨慎设计提示词

本研究探讨前沿大型语言模型在科学馆问答场景中的适用性,旨在提升参观者参与度的同时保证事实准确性。基于英国莱斯特国家太空中心收集的问题数据集,评估了OpenAI的GPT-4、Anthropic的Claude 3.5 Sonnet和Google的Gemini 1.5三个模型的表现。所有模型均被要求生成面向8岁儿童的标准与创意型回答,并由空间科学专家从准确性、吸引力、清晰度、新颖性及偏离预期程度等方面进行评分。结果表明,在保持清晰度和吸引儿童方面,Claude表现最佳,即使在鼓励创意回应时亦然。但总体趋势显示,更高的新颖性通常伴随更低的事实可靠性。研究强调大模型在教育场景中的潜力,同时指出需通过精细提示工程平衡互动性与科学严谨性。

原文摘要 · Abstract (English)

This paper investigates the suitability of frontier Large Language Models (LLMs) for Q&A interactions in science centres, with the aim of boosting visitor engagement while maintaining factual accuracy. Using a dataset of questions collected from the National Space Centre in Leicester (UK), we evaluated responses generated by three leading models: OpenAI's GPT-4, Claude 3.5 Sonnet, and Google Gemini 1.5. Each model was prompted for both standard and creative responses tailored to an 8-year-old audience, and these responses were assessed by space science experts based on accuracy, engagement, clarity, novelty, and deviation from expected answers. The results revealed a trade-off between creativity and accuracy, with Claude outperforming GPT and Gemini in both maintaining clarity and engaging young audiences, even when asked to generate more creative responses. Nonetheless, experts observed that higher novelty was generally associated with reduced factual reliability across all models. This study highlights the potential of LLMs in educational settings, emphasizing the need for careful prompt engineering to balance engagement with scientific rigor.

大模型科学教育问答系统儿童交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。