arXiv:2606.21359cs.CL2026-06

科学微调反增幻觉,18个模型均可靠性下降

Does Finetuning with Scientific Data Increase Hallucinations? A Multi-domain Factuality Evaluation of LLMs

论文配图:Does Finetuning with Scientific Data Increase Hallucinations? A Multi-domain Factuality Evaluation of LLMs
图 1 · 摘自论文原文
  • 构建跨五领域2500条提示的评测基准,区分三类幻觉
  • 微调模型事实错误率上升,虽更自信却更不靠谱
  • 当前查证工具效果有限,专家对可信判断难一致

大语言模型在科学传播中日益普及,但其幻觉风险不容忽视。现有研究多局限于生物医学领域,将幻觉视为二元问题,且未评估日益增多的科学微调模型。本文提出SciFactCheck基准,包含跨五个科学领域的2500个提示,并设计模块化评估框架,针对不可验证、过度宣称和归属错误三类事实性幻觉。通过最小配对控制实验,对比18个科学微调模型与其通用基线模型。结果表明:科学微调模型在所有幻觉类型和领域中事实可靠性均下降;微调模型内部置信度降低,但语言表达更强势。人工预实验显示,当前查证工具与专家判断仅中等一致,甚至人类标注者对“可查证”科学主张也存在分歧。研究挑战了当前领域微调提升事实性的做法,呼吁构建更强的科学内容验证基础设施。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to communicate and explain scientific concepts, yet their tendency to hallucinate poses significant risks in this high stakes use-case. Prior scientific hallucination evaluation work remains largely restricted to the biomedical domain, treats hallucination as a binary task, and has not examined the growing family of scientifically fine-tuned LLMs. We address these gaps with SciFactCheck, a benchmark of 2,500 prompts across five scientific domains, paired with a modular evaluation framework targeting three factuality hallucination types: unverifiability, overclaim, and attribution. Using a controlled minimal-pairing design, we evaluate 18 LLMs by comparing each scientifically fine-tuned model against its general-purpose base. Our results indicate that 1. Scientifically fine-tuned models exhibit degraded factual reliability across all hallucination types and scientific domains, and 2. Fine-tuned models are internally less confident yet linguistically more assertive. A human pilot study further reveals that current fact-checking tools show only modest agreement with expert judgments on scientific content, and that defining scientifically check-worthy claims remains contested even among human annotators. Our findings fundamentally challenge current methods of domain-specific fine-tuning for factuality and call for developing improved verification infrastructure for scientific content.

幻觉检测科学生成模型评估微调风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。