arXiv:2508.17028cs.CL2025-08中稿 · COLM被引 4

用实体搜索提升大模型对表格的理解能力。

Improving Table Understanding with LLMs and Entity-Oriented Search

  • 基于实体的语义搜索,减少预处理和关键词匹配依赖。
  • 在WikiTableQuestions和TabFact上达到新最佳性能。
  • 首次引入图查询语言,拓展表格理解新方向。

本文针对表格理解中的挑战展开研究。现有方法常因表格内容的不可预测性,依赖复杂的预处理和关键词匹配,且缺乏上下文信息,制约了大语言模型(LLMs)的推理能力。为此,我们提出一种面向实体的搜索方法,有效利用问题与表格数据间的语义相似性以及单元格间的隐含关系,减少预处理需求并增强上下文清晰度。该方法聚焦表格实体,使单元格语义紧密关联。此外,我们首创使用图查询语言进行表格理解,开辟了新研究方向。实验表明,该方法在标准基准WikiTableQuestions和TabFact上均达到新的最先进水平。

原文摘要 · Abstract (English)

Our work addresses the challenges of understanding tables. Existing methods often struggle with the unpredictable nature of table content, leading to a reliance on preprocessing and keyword matching. They also face limitations due to the lack of contextual information, which complicates the reasoning processes of large language models (LLMs). To overcome these challenges, we introduce an entity-oriented search method to improve table understanding with LLMs. This approach effectively leverages the semantic similarities between questions and table data, as well as the implicit relationships between table cells, minimizing the need for data preprocessing and keyword matching. Additionally, it focuses on table entities, ensuring that table cells are semantically tightly bound, thereby enhancing contextual clarity. Furthermore, we pioneer the use of a graph query language for table understanding, establishing a new research direction. Experiments show that our approach achieves new state-of-the-art performances on standard benchmarks WikiTableQuestions and TabFact.

表格理解大模型实体搜索图查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。