用人类与AI协作方式,高效构建高质量科学数据集
Using a Human-AI Teaming Approach to Create and Curate Scientific Datasets with the SCILIRE System
- 采用人机协同流程,让研究人员审核并修正AI生成的数据
- 通过反馈机制提升大模型未来推理的准确性
- 在多领域案例中验证了数据提取精度和创建效率
科学文献的快速增长使得手工提取结构化知识变得越来越不切实际。为应对这一挑战,我们提出了SCILIRE系统,用于从科学文献中创建数据集。SCILIRE围绕人机协同原则设计,聚焦于数据验证与清洗的工作流。该系统支持研究人员迭代审查并修正大语言模型输出,同时将这些交互作为反馈信号,用于改进后续的LLM推理。我们通过内在基准测试与多个领域的实际案例研究评估了该设计。结果表明,SCILIRE显著提升了信息提取的准确率,并实现了高效的数据集构建。
原文摘要 · Abstract (English)
The rapid growth of scientific literature has made manual extraction of structured knowledge increasingly impractical. To address this challenge, we introduce SCILIRE, a system for creating datasets from scientific literature. SCILIRE has been designed around Human-AI teaming principles centred on workflows for verifying and curating data. It facilitates an iterative workflow in which researchers can review and correct AI outputs. Furthermore, this interaction is used as a feedback signal to improve future LLM-based inference. We evaluate our design using a combination of intrinsic benchmarking outcomes together with real-world case studies across multiple domains. The results demonstrate that SCILIRE improves extraction fidelity and facilitates efficient dataset creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。