arXiv:2607.24880cs.LGcs.AI2026-07中稿 · IJCAI

让表格嵌入更符合人类对相似性的判断。

Human Preference aligned Tabular Similarity

论文配图:Human Preference aligned Tabular Similarity
图 1 · 摘自论文原文
  • 设计新评估流程,用人类偏好检验表格相似性
  • 发现现有模型在人类判断上表现不佳
  • 适合做企业级数据相似搜索的优化

任务无关的表格嵌入正被广泛应用于产品生命周期管理(PLM)等实际业务系统中的相似性搜索。然而,主流嵌入方法主要针对预测任务优化,未能生成与人类偏好一致的相似性排序。我们指出,传统下游指标无法全面评估嵌入在相似性搜索中的可信度,人类偏好对齐评估是必要且当前缺失的一环。本文提出具体评估流程,并通过PLM场景展示该问题的现实影响。

原文摘要 · Abstract (English)

Task-agnostic tabular embeddings are increasingly used for similarity search in real-world business systems such as Product Lifecycle Management (PLM). However, leading embedding approaches are optimized primarily for prediction tasks - not for producing human preference aligned similarity rankings. We argue that standard downstream metrics are insufficient to fully assess embedding trustworthiness for similarity search and that human preference aligned evaluation is a necessary and currently missing component. We present a concrete evaluation procedure and illustrate the problem through a PLM use case.

表格嵌入相似性搜索人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。