arXiv:2506.12990cs.DBcs.LG2025-06中稿 · ed被引 1

研究人类如何判断表格能否合并,提升数据发现效率。

Humans, Machine Learning, and Language Models in Union: A Cognitive Study on Table Unionability

  • 通过实验分析人类判断表格合并能力的行为模式。
  • 基于观察构建机器学习模型,可提升人类原始判断性能。
  • 发现人类与大模型结合效果优于单独使用,适合人机协同系统设计。

数据发现,特别是表格合并可行性判断,已成为现代数据科学的关键任务。然而,人类在该过程中的行为仍缺乏深入研究。为此,本研究设计并开展了一项实验调查,对人类判断表格合并可行性的决策过程进行了全面分析。基于分析结果,我们构建了一个机器学习框架,用于提升人类的原始判断性能。此外,我们还初步比较了大语言模型(LLM)与人类的表现,发现结合两者通常优于单一使用。本研究为未来高效的人机协同数据发现系统奠定了基础。

原文摘要 · Abstract (English)

Data discovery and table unionability in particular became key tasks in modern Data Science. However, the human perspective for these tasks is still under-explored. Thus, this research investigates the human behavior in determining table unionability within data discovery. We have designed an experimental survey and conducted a comprehensive analysis, in which we assess human decision-making for table unionability. We use the observations from the analysis to develop a machine learning framework to boost the (raw) performance of humans. Furthermore, we perform a preliminary study on how LLM performance is compared to humans indicating that it is typically better to consider a combination of both. We believe that this work lays the foundations for developing future Human-in-the-Loop systems for efficient data discovery.

数据发现人机协同大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。