构建工业级知识图谱问答评估框架,全面衡量系统性能。
Comprehensive Evaluation for a Large Scale Knowledge Graph Question Answering Service
- 设计多层级评估体系,覆盖端到端与组件级指标
- 支持大规模数据集扩展,提前预估系统表现
- 助力真实系统优化,推动数据驱动决策
基于知识图谱的问答系统(KGQA)通过理解自然语言查询中的关系与实体,将其映射为对知识图谱的结构化查询以回答事实类问题。由于系统涉及多个复杂组件,其评估极具挑战性。本文提出 Chronos——一个面向工业规模的 KGQA 系统综合评估框架,旨在实现端到端与组件级指标并重、支持多样数据集扩展,并提供系统上线前的可扩展性能评估方法。我们分析了工业级评估的独特挑战,介绍了 Chronos 的设计思路及其应对方案,展示了该框架如何为数据驱动决策提供基础,并讨论了在真实系统中应用时面临的实际困难。
原文摘要 · Abstract (English)
Question answering systems for knowledge graph (KGQA), answer factoid questions based on the data in the knowledge graph. KGQA systems are complex because the system has to understand the relations and entities in the knowledge-seeking natural language queries and map them to structured queries against the KG to answer them. In this paper, we introduce Chronos, a comprehensive evaluation framework for KGQA at industry scale. It is designed to evaluate such a multi-component system comprehensively, focusing on (1) end-to-end and component-level metrics, (2) scalable to diverse datasets and (3) a scalable approach to measure the performance of the system prior to release. In this paper, we discuss the unique challenges associated with evaluating KGQA systems at industry scale, review the design of Chronos, and how it addresses these challenges. We will demonstrate how it provides a base for data-driven decisions and discuss the challenges of using it to measure and improve a real-world KGQA system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。