用大模型辅助跨学科文献综述,准确率超95%。
Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science
- 用大模型提取论文证据回答研究问题。
- 复现引文准确率超95%,答题准确率约83%。
- 适合需要高效文献梳理的研究者。
大型语言模型(LLMs)被用于协助澳大利亚联邦科学与工业研究组织(CSIRO)的四位研究人员开展系统性文献综述(SLR)。我们评估了这些案例中大模型在文献综述任务中的表现。在每个案例中,我们探讨了参数变化对模型回答准确性的影响。大模型被要求从选定学术论文中提取证据以回答具体研究问题。我们通过专家评审和对比模型答案与专家答案的变换器嵌入余弦相似度,评估模型表现。开发了语义文本高亮工具以支持专家对模型输出的审查。结果显示,当前最先进的大模型在复现文献引文方面准确率超过95%,在回答研究问题方面的准确率约为83%。两种评估方法的相关系数在0.48至0.77之间,表明嵌入余弦相似度可作为衡量语义相似性的有效指标。
原文摘要 · Abstract (English)
Large Language Models (LLMs) were used to assist four Commonwealth Scientific and Industrial Research Organisation (CSIRO) researchers to perform systematic literature reviews (SLR). We evaluate the performance of LLMs for SLR tasks in these case studies. In each, we explore the impact of changing parameters on the accuracy of LLM responses. The LLM was tasked with extracting evidence from chosen academic papers to answer specific research questions. We evaluate the models' performance in faithfully reproducing quotes from the literature and subject experts were asked to assess the model performance in answering the research questions. We developed a semantic text highlighting tool to facilitate expert review of LLM responses. We found that state of the art LLMs were able to reproduce quotes from texts with greater than 95% accuracy and answer research questions with an accuracy of approximately 83%. We use two methods to determine the correctness of LLM responses; expert review and the cosine similarity of transformer embeddings of LLM and expert answers. The correlation between these methods ranged from 0.48 to 0.77, providing evidence that the latter is a valid metric for measuring semantic similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。