arXiv:2511.02936cs.LGcs.AI2025-11

用大模型零样本识别论文中数据的使用方式,省去人工标注

Zero-shot data citation function classification using transformer-based large language models (LLMs)

  • 直接用Llama 3.1-405B模型零样本分类数据引用用途
  • 在无预定义类别下达到F1 0.674,表现优于传统方法
  • 适合做文献数据使用分析的研究者和期刊审稿人

近年来,识别科学文献与特定数据集之间关联的努力不断增多。已知某篇论文引用了某个数据集后,下一步是探究该数据被如何或为何使用。近年来基于预训练的变压器大语言模型(LLMs)的发展,为规模化描述文献中的数据使用场景提供了可能,避免了昂贵的人工标注及经典机器学习系统所需训练数据集的构建。本文应用开源大模型Llama 3.1-405B,对已知包含特定基因组数据集的出版物生成结构化数据使用案例标签。同时引入一种新颖的评估框架以衡量方法有效性。结果表明,未经微调的原始模型在零样本数据引用分类任务中实现了F1分数0.674,尽管前景乐观,但受限于数据可得性、提示词过拟合、计算基础设施及负责任性能评估的成本。

原文摘要 · Abstract (English)

Efforts have increased in recent years to identify associations between specific datasets and the scientific literature that incorporates them. Knowing that a given publication cites a given dataset, the next logical step is to explore how or why that data was used. Advances in recent years with pretrained, transformer-based large language models (LLMs) offer potential means for scaling the description of data use cases in the published literature. This avoids expensive manual labeling and the development of training datasets for classical machine-learning (ML) systems. In this work we apply an open-source LLM, Llama 3.1-405B, to generate structured data use case labels for publications known to incorporate specific genomic datasets. We also introduce a novel evaluation framework for determining the efficacy of our methods. Our results demonstrate that the stock model can achieve an F1 score of .674 on a zero-shot data citation classification task with no previously defined categories. While promising, our results are qualified by barriers related to data availability, prompt overfitting, computational infrastructure, and the expense required to conduct responsible performance evaluation.

大模型零样本数据引用基因组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。