arXiv:2502.15721cs.IRcs.AI2025-02被引 1

用大模型定制生成科研问答数据集,提升知识检索效率。

iTRI-QA: a Toolset for Customized Question-Answer Dataset Generation Using Language Models for Enhanced Scientific Research

  • 结合人工标注与论文库,用微调大模型生成高质量问答对。
  • 构建结构化论文数据库,使回答更贴合科研上下文。
  • 适合需要高效获取文献知识的科研人员快速部署使用。

人工智能在科学领域的迅猛发展亟需高效、可扩展的信息检索与保存方案。本文提出一种名为 iTRI-QA 的工具,用于基于语言模型(LMs)定制生成科研问答(QA)数据集,以支持研究人员以问答形式高效获取科学知识。该方法将经过筛选的问答数据集与专用科研论文数据集相结合,通过微调语言模型提升回答的上下文相关性与准确性。整个流程包含四个关键步骤:(1) 生成高质量的人工标注问答样本;(2) 构建结构化的科研论文数据库;(3) 使用领域特定的问答样本对语言模型进行微调;(4) 生成符合用户查询并匹配已有的精选数据库的问答数据集。该系统具备动态性和领域专属性,显著增强语言模型在学术研究中的应用价值,为未来科研语言模型的部署奠定基础。我们验证了该工具在科学知识检索场景下的可行性与可扩展性,为其在跨学科应用中的集成铺平道路。

原文摘要 · Abstract (English)

The exponential growth of AI in science necessitates efficient and scalable solutions for retrieving and preserving research information. Here, we present a tool for the development of a customized question-answer (QA) dataset, called Interactive Trained Research Innovator (iTRI) - QA, tailored for the needs of researchers leveraging language models (LMs) to retrieve scientific knowledge in a QA format. Our approach integrates curated QA datasets with a specialized research paper dataset to enhance responses' contextual relevance and accuracy using fine-tuned LM. The framework comprises four key steps: (1) the generation of high-quality and human-generated QA examples, (2) the creation of a structured research paper database, (3) the fine-tuning of LMs using domain-specific QA examples, and (4) the generation of QA dataset that align with user queries and the curated database. This pipeline provides a dynamic and domain-specific QA system that augments the utility of LMs in academic research that will be applied for future research LM deployment. We demonstrate the feasibility and scalability of our tool for streamlining knowledge retrieval in scientific contexts, paving the way for its integration into broader multi-disciplinary applications.

科研问答大模型知识检索数据集生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。