arXiv:2512.22396cs.AIcond-mat.mtrl-sci2025-12被引 4

用多阶段验证减少大模型材料科学内容的幻觉问题。

HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification

  • 分四步验证:内在一致性、多源检索、矛盾图分析、指标评估。
  • 幻觉率降低30%,高熵问题更易出错。
  • 新指标可测语义等价查询间的不一致,适合科研可信度评估。

人工智能特别是大语言模型(LLMs)正在改变科学发现,加速知识生成与假设提出。然而,幻觉问题——即模型生成事实错误或误导性信息——严重威胁研究可靠性。为此,我们构建了HalluMatData基准数据集,用于评估幻觉检测、事实一致性和响应鲁棒性。同时提出HalluMatDetector多阶段检测框架,融合内在验证、多源检索、矛盾图分析与指标评估,有效识别并缓解幻觉。结果表明,不同材料科学子领域幻觉程度差异显著,高熵查询更易出现事实不一致。使用该框架后,幻觉率相比标准输出降低30%。此外,我们引入改写式幻觉一致性评分(PHCS),量化语义等价查询中模型响应的不一致性,为评估模型可靠性提供新视角。

原文摘要 · Abstract (English)

Artificial Intelligence (AI), particularly Large Language Models (LLMs), is transforming scientific discovery, enabling rapid knowledge generation and hypothesis formulation. However, a critical challenge is hallucination, where LLMs generate factually incorrect or misleading information, compromising research integrity. To address this, we introduce HalluMatData, a benchmark dataset for evaluating hallucination detection methods, factual consistency, and response robustness in AI-generated materials science content. Alongside this, we propose HalluMatDetector, a multi-stage hallucination detection framework that integrates intrinsic verification, multi-source retrieval, contradiction graph analysis, and metric-based assessment to detect and mitigate LLM hallucinations. Our findings reveal that hallucination levels vary significantly across materials science subdomains, with high-entropy queries exhibiting greater factual inconsistencies. By utilizing HalluMatDetector verification pipeline, we reduce hallucination rates by 30% compared to standard LLM outputs. Furthermore, we introduce the Paraphrased Hallucination Consistency Score (PHCS) to quantify inconsistencies in LLM responses across semantically equivalent queries, offering deeper insights into model reliability.

幻觉检测材料科学大模型验证框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。