arXiv:2505.12495cs.CL2025-05被引 2

基于知识图谱的多层级问答抽取框架,可精准评估大模型长文本推理能力。

KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation

  • 用知识图谱表示文档,按多跳检索、集合运算、答案多样性分层抽取问答对
  • 构建20,139个金融信贷协议问答对,发现顶尖模型在集合比较和多跳推理上仍表现不佳
  • 揭示模型因语义误解和隐含关系处理失败导致系统性错误,适合评估长文本推理模型

我们提出KG-MuLQA(基于知识图谱的多层级问答抽取):一个框架,能够(1)在多个复杂度层级上提取问答对;(2)沿三个关键维度——多跳检索、集合运算和答案多样性——进行抽取;(3)通过知识图谱驱动的文档表示实现。该方法支持在受控难度下对模型性能进行细粒度评估。利用此框架,我们在金融信贷协议基础上构建了包含20,139个问答对的数据集,并评估了16个专有及开源大语言模型。结果表明,即使表现最佳的模型在集合比较和长上下文多跳推理方面仍存在困难。分析揭示了与语义误读及隐含关系处理能力不足相关的系统性失败模式。

原文摘要 · Abstract (English)

We introduce KG-MuLQA (Knowledge-Graph-based Multi-Level Question-Answer Extraction): a framework that (1) extracts QA pairs at multiple complexity levels (2) along three key dimensions -- multi-hop retrieval, set operations, and answer plurality, (3) by leveraging knowledge-graph-based document representations. This approach enables fine-grained assessment of model performance across controlled difficulty levels. Using this framework, we construct a dataset of 20,139 QA pairs based on financial credit agreements and evaluate 16 proprietary and open-weight Large Language Models, observing that even the best-performing models struggle with set-based comparisons and multi-hop reasoning over long contexts. Our analysis reveals systematic failure modes tied to semantic misinterpretation and inability to handle implicit relations.

知识图谱问答系统长文本推理多跳问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。