arXiv:2507.15715cs.CLastro-ph.IM2025-07中稿 · COLM被引 4

通过分析天文学家使用AI工具的提问与评价,改进科学领域大模型评估方法。

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs

  • 基于368条真实查询和11位天文学家访谈,提炼出用户评估标准。
  • 发现用户关注回答准确性、上下文相关性及对专业术语的正确理解。
  • 提出可复用的评估框架,适合科研场景的大模型评测。

随着大语言模型(LLMs)在天文学等科研领域的应用日益广泛,现有通用评估基准未能跟上用户多样化的真实使用方式。本研究聚焦一个具体应用场景:通过Slack部署的检索增强生成型天文文献问答机器人。通过对四周期间368条用户查询进行归纳编码,并对11位天文学家开展后续访谈,揭示了人类用户在评估该系统时所使用的提问类型与判断标准。研究将发现整合为具体建议,用于构建面向天文学领域的样本评估基准。整体工作为提升大模型在科学研究中的评估质量与实际可用性提供了有效路径。

原文摘要 · Abstract (English)

There is growing interest in leveraging LLMs to aid in astronomy and other scientific research, but benchmarks for LLM evaluation in general have not kept pace with the increasingly diverse ways that real people evaluate and use these models. In this study, we seek to improve evaluation procedures by building an understanding of how users evaluate LLMs. We focus on a particular use case: an LLM-powered retrieval-augmented generation bot for engaging with astronomical literature, which we deployed via Slack. Our inductive coding of 368 queries to the bot over four weeks and our follow-up interviews with 11 astronomers reveal how humans evaluated this system, including the types of questions asked and the criteria for judging responses. We synthesize our findings into concrete recommendations for building better benchmarks, which we then employ in constructing a sample benchmark for evaluating LLMs for astronomy. Overall, our work offers ways to improve LLM evaluation and ultimately usability, particularly for use in scientific research.

大模型评估天文学人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。