arXiv:2506.06371cs.CL2025-06被引 1

结合统计分析与大模型,高效识别表格列间关系。

Relationship Detection on Tabular Data Using Statistical Analysis and Large Language Models

  • 用统计方法缩小大模型搜索范围,提升效率
  • 在两个基准数据集上达到领先水平
  • 适合做表格语义理解的研究者和工程师

近年来,由于其重要性及新技术和基准的引入,表格解析任务取得了显著进展。本文提出一种混合方法,用于检测未标注表格中列之间的关系(即CPA任务),以知识图谱(KG)为参考。该方法利用大语言模型(LLMs)并结合统计分析,减少潜在KG关系的搜索空间。主要模块包括领域与范围约束检测、关系共现分析。在SemTab挑战提供的两个基准数据集上进行实验,评估了各模块的影响以及不同先进LLM在多种量化级别下的表现。实验还对比了不同提示策略。所提方法已在GitHub公开,结果表明其在这些数据集上具有竞争力。

原文摘要 · Abstract (English)

Over the past few years, table interpretation tasks have made significant progress due to their importance and the introduction of new technologies and benchmarks in the field. This work experiments with a hybrid approach for detecting relationships among columns of unlabeled tabular data, using a Knowledge Graph (KG) as a reference point, a task known as CPA. This approach leverages large language models (LLMs) while employing statistical analysis to reduce the search space of potential KG relations. The main modules of this approach for reducing the search space are domain and range constraints detection, as well as relation co-appearance analysis. The experimental evaluation on two benchmark datasets provided by the SemTab challenge assesses the influence of each module and the effectiveness of different state-of-the-art LLMs at various levels of quantization. The experiments were performed, as well as at different prompting techniques. The proposed methodology, which is publicly available on github, proved to be competitive with state-of-the-art approaches on these datasets.

表格理解大模型关系检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。