提出可解释的开放域事实一致性评估框架,解决大模型幻觉问题
AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
- 将文本分解为基本事实,采用无模式灵活评估方法
- 引入加权指标提升事实评估准确性,支持复杂领域控制评估难度
- 在通用与临床数据集上验证有效,代码开源供后续研究使用
大型语言模型在自然语言处理任务中取得显著进展,但仍容易生成看似合理却错误或误导性的内容,这一现象被称为幻觉,在临床等高风险领域尤为严重。现有评估指标无法充分衡量事实一致性且缺乏可解释性,导致错误诊断和修正困难。为此,我们提出一种针对领域内与开放域文本的事实一致性评估可解释框架。该方法将文本分解为原子事实,采用灵活的无模式评估策略;不同于以往绝对度量方式,引入加权指标以增强事实评估能力;并设计机制控制复杂领域的评估复杂度。我们在多个主流通用及临床数据集上进行了基准测试,并开源代码,以支持未来事实感知模型的训练研究。
原文摘要 · Abstract (English)
Large Language Models have significantly advanced natural language processing tasks, but remain prone to generating incorrect or misleading but plausible arguments. This issue, known as hallucination, is particularly concerning in high-stakes domains like clinical applications, where factual inaccuracies can have severe consequences. Existing evaluation metrics fail to adequately assess factual consistency and lack interpretability, making diagnosing and mitigating errors difficult. We propose an interpretable framework for factual consistency assessment for in-domain and open-domain texts to address these limitations. Our approach decomposes text into atomic facts and introduces a flexible, schema-free methodology. Unlike previous methods with an absolute metric, we incorporate a weighted metric to enhance factual evaluation. Additionally, we propose a mechanism to control assessment complexity in intricate domains. We benchmark our approach on popular general and clinical datasets and release our code to support fact-aware model training in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。