研究大模型如何因内部知识误判真实证据,导致生成错误流程图。
Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- 设计冲突场景测试模型对输入证据的忠实度
- 标准流程下模型仍常偏离真实描述
- 适合关注AI可靠性与验证的从业者
大语言模型在分析任务中的价值源于其海量预训练知识,可解读模糊输入并补全缺失信息。然而,这一能力也带来关键风险:知识驱动型幻觉——模型输出与显式源证据矛盾,因内部通用知识覆盖了真实信息。本文通过自动化流程建模任务研究该现象,目标是从源文档生成正式业务流程模型。业务流程管理(BPM)领域提供理想研究场景,因多数核心流程遵循标准化模式,模型具备强预训练模式。我们设计受控实验,制造输入证据与模型背景知识间的故意冲突,使用描述标准与刻意异常流程结构的输入,测量模型对提供证据的忠实度。研究提出评估该可靠性问题的方法论,警示任何基于证据的领域需严格验证AI生成结果。
原文摘要 · Abstract (English)
The utility of Large Language Models (LLMs) in analytical tasks is rooted in their vast pre-trained knowledge, which allows them to interpret ambiguous inputs and infer missing information. However, this same capability introduces a critical risk of what we term knowledge-driven hallucination: a phenomenon where the model's output contradicts explicit source evidence because it is overridden by the model's generalized internal knowledge. This paper investigates this phenomenon by evaluating LLMs on the task of automated process modeling, where the goal is to generate a formal business process model from a given source artifact. The domain of Business Process Management (BPM) provides an ideal context for this study, as many core business processes follow standardized patterns, making it likely that LLMs possess strong pre-trained schemas for them. We conduct a controlled experiment designed to create scenarios with deliberate conflict between provided evidence and the LLM's background knowledge. We use inputs describing both standard and deliberately atypical process structures to measure the LLM's fidelity to the provided evidence. Our work provides a methodology for assessing this critical reliability issue and raises awareness of the need for rigorous validation of AI-generated artifacts in any evidence-based domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。