用知识图谱评估大模型的上下文理解能力,更准地判断回答是否靠谱。
Evaluation of Contextual Understanding in Large Language Models

- 基于知识图谱设计结构与语义融合的相似度评分
- 在问答数据集上验证了新方法对正确性和忠实性的检测能力
- 适合研究模型推理可靠性或开发可解释AI的人看
大语言模型在多种自然语言任务中表现优异,但其真正理解上下文的能力仍存疑。传统评估指标如困惑度、BLEU或表层准确率无法揭示模型提取、整合和推理上下文信息的能力——这一缺陷在问答任务中尤为关键,因模型需基于上下文知识作答,而非依赖记忆关联。本文提出一种基于知识图谱的新型评估框架,引入S3KG(语义结构相似性)作为混合相似度度量,结合连续评分机制,并构建诊断框架以分类推理错误。通过在精心构建的问答基准上评估,证明该方法在衡量生成回答的正确性、忠实性与可解释性方面优于现有指标。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs extract, integrate, and reason over contextual information--a gap particularly critical in question answering, where models must align responses with contextually grounded knowledge rather than memorized associations. We propose a novel knowledge graph-based evaluation framework introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors. To validate this pipeline, we evaluate S3KG against established metrics on a curated question-answer (QA) benchmark, demonstrating its effectiveness in measuring correctness, faithfulness, and interpretability in LLM-generated responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。