解决大模型在科学文献中误报测量数据的问题。
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

- 构建细粒度测量幻觉分类体系,识别量、单位、修饰词等错误。
- 通过两阶段推理微调与渐进奖励机制,降低幻觉率32%以上。
- 适合需要高精度科学数据提取的研究者与自动化分析系统。
从科学文献中准确提取测量数据是AI4Science中的关键挑战,但大语言模型常出现严重幻觉,严重影响自动文献理解系统的可靠性。为此,我们提出MeasHalu框架,通过增强推理与针对性优化缓解科学测量幻觉。首先构建细粒度的测量幻觉分类体系,涵盖量、单位、修饰词及关系等维度;采用基于增强科学数据与过程监督的两阶段推理感知微调策略;引入渐进式奖励课程,对特定幻觉类型施加惩罚,显著提升提取忠实度。实验表明,MeasHalu在MeasEval基准上大幅降低幻觉率并提高整体准确率。本工作为自动化科学知识提取中的核心瓶颈提供了针对性解决方案,推动更可信、可扩展的机器辅助科学文献分析。
原文摘要 · Abstract (English)
The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and integration of quantitative research findings. However, Large Language Models (LLMs) frequently exhibit severe hallucinations, which significantly undermine the reliability of automated scientific document understanding systems. To address this problem, we propose MeasHalu, a novel framework for mitigating scientific measurement hallucinations through enhanced reasoning and targeted optimization. We first present a fine-grained taxonomy of measurement-specific hallucinations, categorizing errors across quantities, units, modifiers, and relations. Our approach incorporates a two-stage reasoning-aware fine-tuning strategy using augmented scientific data and process-based supervision. Furthermore, we introduce a progressive reward curriculum designed to penalize specific hallucination types, significantly improving extraction faithfulness. Experimental results demonstrate that MeasHalu substantially reduces hallucination rates and improves overall accuracy on the MeasEval benchmark. This work provides a targeted solution to a key bottleneck in automated scientific knowledge extraction, facilitating more trustworthy and scalable machine-assisted scientific literature analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。