arXiv:2511.06973cs.LGcs.CV2025-11中稿 · EurIPS'25: AI for …

用空间布局和数据类型打造新相似度度量,精准识别电子表格模板。

Oh That Looks Familiar: A Novel Similarity Measure for Spreadsheet Template Discovery

  • 融合语义嵌入、数据类型与位置信息,构建细胞级相似度计算
  • 在FUSTE数据集上实现1.00的调整兰德指数,完美还原模板结构
  • 适合需要批量发现模板的自动化数据处理场景

传统方法难以捕捉电子表格的结构布局和类型模式。为此,我们提出一种混合距离度量,结合语义嵌入、数据类型与空间位置信息。通过将电子表格转换为细胞级嵌入,并使用Chamfer和Hausdorff距离等聚合技术计算相似性。在多个模板族上的实验表明,相比基于图的Mondrian基线,本方法在无监督聚类中表现更优,在FUSTE数据集上达到1.00的调整兰德指数(对比0.90),实现完美模板重建。该方法支持大规模自动化模板发现,可应用于表格集合的检索增强生成、模型训练及批量数据清洗等下游任务。

原文摘要 · Abstract (English)

Traditional methods for identifying structurally similar spreadsheets fail to capture the spatial layouts and type patterns defining templates. To quantify spreadsheet similarity, we introduce a hybrid distance metric that combines semantic embeddings, data type information, and spatial positioning. In order to calculate spreadsheet similarity, our method converts spreadsheets into cell-level embeddings and then uses aggregation techniques like Chamfer and Hausdorff distances. Experiments across template families demonstrate superior unsupervised clustering performance compared to the graph-based Mondrian baseline, achieving perfect template reconstruction (Adjusted Rand Index of 1.00 versus 0.90) on the FUSTE dataset. Our approach facilitates large-scale automated template discovery, which in turn enables downstream applications such as retrieval-augmented generation over tabular collections, model training, and bulk data cleaning.

表格模板相似度度量自动化处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。