arXiv:2411.03356cs.LGcs.AI2024-11被引 7

用大模型生成表格相似性数据,提升表格推荐效果

Enhancing Table Representations with LLM-powered Synthetic Data Generation

  • 利用LLM自动生成符合表格相似性的合成数据
  • 合成数据使表格推荐准确率显著提升
  • 适合做表格理解与推荐系统的研究者

在数据驱动决策时代,精准的表格级表示与高效的表格推荐系统对表格管理、发现与分析日益重要。然而,现有表格表示方法多聚焦单元格级任务,且缺乏高质量训练数据。为此,我们首先定义了企业数据转换场景下的表格相似性标准,作为合成数据生成的基础。基于此,提出一种新型合成数据生成流程,利用大语言模型(LLMs)的代码生成与数据操作能力,构建大规模面向表格级表示学习的合成数据集。通过人工验证及表格推荐任务上的性能对比,证明该流程生成的合成数据符合预设的表格相似性定义,显著增强表格表示,进而提升推荐性能。

原文摘要 · Abstract (English)

In the era of data-driven decision-making, accurate table-level representations and efficient table recommendation systems are becoming increasingly crucial for improving table management, discovery, and analysis. However, existing approaches to tabular data representation often face limitations, primarily due to their focus on cell-level tasks and the lack of high-quality training data. To address these challenges, we first formulate a clear definition of table similarity in the context of data transformation activities within data-driven enterprises. This definition serves as the foundation for synthetic data generation, which require a well-defined data generation process. Building on this, we propose a novel synthetic data generation pipeline that harnesses the code generation and data manipulation capabilities of Large Language Models (LLMs) to create a large-scale synthetic dataset tailored for table-level representation learning. Through manual validation and performance comparisons on the table recommendation task, we demonstrate that the synthetic data generated by our pipeline aligns with our proposed definition of table similarity and significantly enhances table representations, leading to improved recommendation performance.

表格表示合成数据LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。