从155篇化学预印本中提取952个高质量问答对,助力化学NLP研究。
ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers
- 通过OCR+GPT-4o+模糊搜索自动构建问答数据集
- 覆盖17个化学子领域,含概念、机制、实验等类型问题
- 适合做化学领域大模型微调或智能检索系统开发
化学文献的快速扩张给研究人员高效获取专业知识带来挑战。为推动化学领域的自然语言处理发展,我们提出了ChemQuests,一个从155篇ChemRxiv论文中提取的952个高质量问答对数据集,覆盖17个化学子领域。每个问答对均明确关联原文段落,确保可追溯性和上下文准确性。数据集通过自动化流程构建:结合光学字符识别(OCR)、GPT-4o生成问答对,并用模糊搜索进行验证。该数据集侧重概念性、机理性、应用性以及合成或实验类问题,适用于基于检索的问答系统、搜索引擎开发及领域适配的大语言模型微调。我们分析了数据集的结构、覆盖范围与局限性,并提出未来扩展与专家验证方向。ChemQuests为化学NLP研究、教育及工具开发提供了基础资源。
原文摘要 · Abstract (English)
The rapid expansion of chemistry literature poses significant challenges for researchers seeking to efficiently access domain-specific knowledge. To support advancements in chemistry-focused natural language processing (NLP), we present ChemQuests, a curated dataset of 952 high-quality question-answer (QA) pairs derived from 155 ChemRxiv \cite{chemrxivWebsite} papers across 17 subfields of chemistry. Each QA pair is explicitly linked to its source text segment to ensure traceability and contextual accuracy. ChemQuests was constructed using an automated pipeline that combines optical character recognition (OCR), QA generation using GPT-4o, and fuzzy-search verification. The dataset emphasizes conceptual, mechanistic, applied, and synthetic or experimental questions, enabling applications in retrieval-based QA systems, search engine development, and fine-tuning of domain-adapted large language models. We analyze the dataset's structure, coverage, and limitations, and outline future directions for expansion and expert validation. ChemQuests provides a foundational resource for chemistry NLP research, education, and tool development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。