用逻辑链补全提升科学数据集质量,错误率从20%降至2%以下
LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- 通过补全答案中缺失的逻辑步骤并分离科学原理与推导
- 在挑战性数据集上将错误率从最高20%降低至2%以下
- 适合需要高质量科学数据的AI模型训练与评估者
尽管大语言模型在通用领域表现优异,但在科学问题求解中可靠性不足。科学AI的发展依赖大规模高质量语料库。然而现有科学问答数据集存在高错误率,常由答案中的逻辑跳跃和隐含推理导致。为此,我们提出LOCA(逻辑链增强)框架,通过增补-审查循环实现科学语料自动清洗。核心是补全原始答案中缺失的逻辑步骤,并显式区分科学原理与其后续推导。在多个挑战性科学语料库上应用LOCA,可自动过滤噪声数据,通常将错误率从高达20%降至2%以下。该方法为构建高质量科学语料库提供了可扩展且高效的新路径,推动更可靠的科学AI训练与评估。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) excel in general domains, their reliability often falls short in scientific problem-solving. The advancement of scientific AI depends on large-scale, high-quality corpora. However, existing scientific question-answering (QA) datasets suffer from high error rates, frequently resulting from logical leaps and implicit reasoning within the answers. To address this issue, we introduce LOCA (Logical Chain Augmentation), a novel framework for automatically cleaning scientific corpora, implemented through an augment-and-review loop. At its core, LOCA enhances raw answers by completing missing logical steps and explicitly separating the underlying scientific principle from its subsequent derivation. By applying LOCA to challenging scientific corpora, we demonstrate that it can automatically filter noisy datasets, typically reducing the error rate from as high as 20\% to below 2\%. LOCA provides a scalable and effective methodology for creating high-quality scientific corpora, paving the way for more reliable training and evaluation of scientific AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。