arXiv:2502.15718cs.IRcs.LG2025-02被引 2

自动化分析海量公开数据,提升数据发现与利用效率。

Making Sense of Data in the Wild: Data Analysis Automation at Scale

  • 用智能代理+检索增强生成自动分析数据集。
  • 提升数据检索命中率与多样性,报告可改进合成数据质量。
  • 适合数据科学家、ML研究者快速挖掘可用数据资源。

随着公开数据量持续增长,研究人员面临基准任务数据多样性不足的挑战。尽管公共仓库中有数千个数据集,但数量庞大反而使筛选合适数据变得困难,导致许多优质数据未被充分利用。同时,尽管长期倡导提升数据整理质量,现有方法仍耗时耗力。本文提出一种新方法,结合智能代理与检索增强生成,实现大规模数据自动分析、数据集整理与索引。系统通过多个代理分析公共仓库中的原始非结构化数据,生成数据集报告与可交互的可视化索引,便于探索。实验表明,该方法显著提升了数据描述的详尽程度、检索命中率及多样性。此外,生成的数据集报告可被其他机器学习模型使用,提升合成数据生成的准确性和真实感。本方法有效简化了将原始数据转化为可训练数据集的流程,帮助研究者更好利用现有数据资源。

原文摘要 · Abstract (English)

As the volume of publicly available data continues to grow, researchers face the challenge of limited diversity in benchmarking machine learning tasks. Although thousands of datasets are available in public repositories, the sheer abundance often complicates the search for suitable data, leaving many valuable datasets underexplored. This situation is further amplified by the fact that, despite longstanding advocacy for improving data curation quality, current solutions remain prohibitively time-consuming and resource-intensive. In this paper, we propose a novel approach that combines intelligent agents with retrieval augmented generation to automate data analysis, dataset curation and indexing at scale. Our system leverages multiple agents to analyze raw, unstructured data across public repositories, generating dataset reports and interactive visual indexes that can be easily explored. We demonstrate that our approach results in more detailed dataset descriptions, higher hit rates and greater diversity in dataset retrieval tasks. Additionally, we show that the dataset reports generated by our method can be leveraged by other machine learning models to improve the performance on specific tasks, such as improving the accuracy and realism of synthetic data generation. By streamlining the process of transforming raw data into machine-learning-ready datasets, our approach enables researchers to better utilize existing data resources.

数据自动化智能代理数据挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。