现有表格合并评测基准存在漏洞,简单方法也能胜过复杂模型。
Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks
- 发现主流评测基准依赖数据集特性,而非真正语义理解。
- 简单基线在多个基准上表现优于复杂模型,最高得分超30%。
- 提出新评测标准,适合研究语义表格合并的学者参考。
近期表格表示学习与数据发现方法致力于解决数据湖中的表格合并搜索(TUS)问题,即找出可与查询表合并以丰富内容的表格。这些方法通常使用旨在评估真实世界TUS任务中语义理解能力的基准进行评测。然而,我们对主流TUS基准的分析揭示了若干缺陷,使得简单基线在其中表现异常出色,常超过更复杂的模型。这表明当前基准分数高度受数据集特性的干扰,无法有效分离出语义理解带来的性能提升。为此,我们提出了未来基准应具备的关键标准,以实现对语义表格合并进展更真实、可靠的评估。
原文摘要 · Abstract (English)
Recent table representation learning and data discovery methods tackle table union search (TUS) within data lakes, which involves identifying tables that can be unioned with a given query table to enrich its content. These methods are commonly evaluated using benchmarks that aim to assess semantic understanding in real-world TUS tasks. However, our analysis of prominent TUS benchmarks reveals several limitations that allow simple baselines to perform surprisingly well, often outperforming more sophisticated approaches. This suggests that current benchmark scores are heavily influenced by dataset-specific characteristics and fail to effectively isolate the gains from semantic understanding. To address this, we propose essential criteria for future benchmarks to enable a more realistic and reliable evaluation of progress in semantic table union search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。