构建法律式评估框架,检测大模型推理中的认知幻觉。
CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
- 借鉴法律证据标准,分级评估模型推理的忠实性
- 发现87%的认知陈述存在幻觉,显著高于事实类幻觉
- 支持自动标注,适合研究幻觉检测与模型评测的团队
大语言模型(LLM)的忠实性幻觉指其生成内容未得到输入上下文支持。现有基准多聚焦于重述型事实陈述,忽视了需从上下文中进行推断的认知类陈述,导致此类幻觉难以评估与检测。受法律领域证据审查启发,我们设计了一套严谨的评估框架,用于衡量不同层级的认知忠实性,并构建了CogniBench数据集,揭示关键统计特征。为适应快速演进的LLM,我们进一步开发了可扩展的自动化标注流水线,形成大规模的CogniBench-L数据集,助力训练精准的事实与认知幻觉检测器。相关模型与数据集已开源:https://github.com/FUTUREEEEEE/CogniBench。
原文摘要 · Abstract (English)
Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on "factual statements" that rephrase source materials while overlooking "cognitive statements" that involve making inferences from the given context. Consequently, evaluating and detecting the hallucination of cognitive statements remains challenging. Inspired by how evidence is assessed in the legal domain, we design a rigorous framework to assess different levels of faithfulness of cognitive statements and introduce the CogniBench dataset where we reveal insightful statistics. To keep pace with rapidly evolving LLMs, we further develop an automatic annotation pipeline that scales easily across different models. This results in a large-scale CogniBench-L dataset, which facilitates training accurate detectors for both factual and cognitive hallucinations. We release our model and datasets at: https://github.com/FUTUREEEEEE/CogniBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。