arXiv:2603.20252cs.CLq-fin.CP2026-03被引 1

构建金融问答幻觉检测基准,评估知识图谱增强系统可靠性。

FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems

  • 基于美国证券交易委员会10-K文件构建带标注的幻觉检测数据集。
  • 嵌入方法在噪声三元组下仍保持9%性能下降,优于其他方法。
  • 为金融、医疗等高风险领域提供AI可信性评估框架。

随着组织将AI问答系统用于合规、风险评估和决策支持,确保生成内容的事实准确性成为关键工程挑战。现有知识图谱(KG)增强型问答系统缺乏系统性幻觉检测机制——即事实错误输出,会损害可靠性并削弱用户信任。本文提出FinBench-QA-Hallucination,一个针对证券10-K文件的金融问答幻觉检测基准。数据集包含755个标注样本,来自300页文档,每条均通过严格证据链协议标注,要求文本段落与提取的关系三元组同时支持。评估六种检测方法:大模型裁判、微调分类器、自然语言推理模型、跨度检测器及基于嵌入的方法,在有无知识图谱三元组两种条件下进行测试。结果显示,大模型裁判与嵌入方法在干净条件下表现最佳(F1: 0.82–0.86);但引入噪声三元组后,多数方法性能显著下降,马修斯相关系数(MCC)下降44%-84%,而嵌入方法仅下降9%。统计检验(Cochran's Q 和 McNemar)证实差异显著(p < 0.001)。研究揭示当前KG增强系统的脆弱性,为构建可靠金融信息系统提供依据,因幻觉可能导致监管违规与错误决策。该基准也为医疗、法律、政府等高风险领域提供可复用的AI可靠性评估框架。

原文摘要 · Abstract (English)

As organizations increasingly integrate AI-powered question-answering systems into financial information systems for compliance, risk assessment, and decision support, ensuring the factual accuracy of AI-generated outputs becomes a critical engineering challenge. Current Knowledge Graph (KG)-augmented QA systems lack systematic mechanisms to detect hallucinations - factually incorrect outputs that undermine reliability and user trust. We introduce FinBench-QA-Hallucination, a benchmark for evaluating hallucination detection methods in KG-augmented financial QA over SEC 10-K filings. The dataset contains 755 annotated examples from 300 pages, each labeled for groundedness using a conservative evidence-linkage protocol requiring support from both textual chunks and extracted relational triplets. We evaluate six detection approaches - LLM judges, fine-tuned classifiers, Natural Language Inference (NLI) models, span detectors, and embedding-based methods under two conditions: with and without KG triplets. Results show that LLM-based judges and embedding approaches achieve the highest performance (F1: 0.82-0.86) under clean conditions. However, most methods degrade significantly when noisy triplets are introduced, with Matthews Correlation Coefficient (MCC) dropping 44-84 percent, while embedding methods remain relatively robust with only 9 percent degradation. Statistical tests (Cochran's Q and McNemar) confirm significant performance differences (p < 0.001). Our findings highlight vulnerabilities in current KG-augmented systems and provide insights for building reliable financial information systems, where hallucinations can lead to regulatory violations and flawed decisions. The benchmark also offers a framework for integrating AI reliability evaluation into information system design across other high-stakes domains such as healthcare, legal, and government.

金融AI幻觉检测知识图谱评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。