arXiv:2411.11891cs.AIcs.IR2024-11综述被引 9

梳理表格语义解析的挑战与方向,助你选对工具

Survey on Semantic Interpretation of Tabular Data: Challenges and Directions

  • 按31个属性分类主流解析方法,方便对比选择
  • 评估12项指标,分析现有工具优劣与适用场景
  • 给出实用指南,适合构建知识图谱或问答系统者参考

表格数据在多个领域中具有关键作用,是数据处理与交换的常见形式,尤其在网页上广泛应用。对表格信息的语义解读、提取与处理对知识密集型应用至关重要。近年来,大量工作致力于将表格数据与背景知识图谱中的本体和实体进行标注,这一过程称为语义表格解析(Semantic Table Interpretation, STI)。STI自动化有助于构建知识图谱、丰富数据内容并提升基于网络的问答能力。本综述旨在全面梳理STI研究现状,首先基于31个属性对现有方法进行分类,便于比较与评估;其次分析可用工具,并依据12个标准进行评价;进一步深入探讨用于评估STI方法的黄金标准;最后为终端用户提供实际指导,帮助其根据任务需求选择合适方法,并讨论未解决的问题及未来可能的研究方向。

原文摘要 · Abstract (English)

Tabular data plays a pivotal role in various fields, making it a popular format for data manipulation and exchange, particularly on the web. The interpretation, extraction, and processing of tabular information are invaluable for knowledge-intensive applications. Notably, significant efforts have been invested in annotating tabular data with ontologies and entities from background knowledge graphs, a process known as Semantic Table Interpretation (STI). STI automation aids in building knowledge graphs, enriching data, and enhancing web-based question answering. This survey aims to provide a comprehensive overview of the STI landscape. It starts by categorizing approaches using a taxonomy of 31 attributes, allowing for comparisons and evaluations. It also examines available tools, assessing them based on 12 criteria. Furthermore, the survey offers an in-depth analysis of the Gold Standards used for evaluating STI approaches. Finally, it provides practical guidance to help end-users choose the most suitable approach for their specific tasks while also discussing unresolved issues and suggesting potential future research directions.

表格解析知识图谱综述语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。