arXiv:2509.25107cs.CL2025-09Conference of the …被引 3

LLMs虽强,但知识抽取仍能提升问答效果

Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?

  • 在现有基准上添加三元组标注,评估不同规模LLM性能
  • 大模型问答准确率高,但仍需三元组增强效果
  • 适合想提升问答系统性能的研究者与工程师

大型语言模型(LLMs)的兴起显著提升了基于半结构化内容的网络问答系统性能,引发对知识抽取是否仍具价值的质疑。本文通过在现有基准中引入知识抽取标注,评估了多种商业与开源、不同规模的LLM。结果表明,网络级知识抽取对LLMs而言仍是挑战。尽管在问答任务中达到较高准确率,但通过引入提取的三元组或采用多任务学习,仍可进一步提升性能。研究揭示了知识三元组抽取在新型问答系统中的新角色,并为不同模型规模与资源条件下最大化LLM效能提供了策略。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLMs) has significantly advanced web-based Question Answering (QA) systems over semi-structured content, raising questions about the continued utility of knowledge extraction for question answering. This paper investigates the value of triple extraction in this new paradigm by extending an existing benchmark with knowledge extraction annotations and evaluating commercial and open-source LLMs of varying sizes. Our results show that web-scale knowledge extraction remains a challenging task for LLMs. Despite achieving high QA accuracy, LLMs can still benefit from knowledge extraction, through augmentation with extracted triples and multi-task learning. These findings provide insights into the evolving role of knowledge triple extraction in web-based QA and highlight strategies for maximizing LLM effectiveness across different model sizes and resource settings.

知识抽取问答系统LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。