arXiv:2412.02211cs.LG2024-12被引 20

用自编码器自动提取特征并降维,效果优于传统方法。

An Automated Data Mining Framework Using Autoencoders for Feature Extraction and Dimensionality Reduction

  • 用自编码器的编码解码结构捕捉数据深层特征。
  • 重建误差和均方根误差更低,更好保留数据结构。
  • 适合希望减少人工干预、提升自动化程度的从业者。

本研究提出一种基于自编码器的自动化数据挖掘框架,并实验验证其在特征提取与降维方面的有效性。通过编码-解码结构,自编码器能捕捉数据潜在特征,实现降噪与异常检测,为数据挖掘提供高效稳定的解决方案。实验对比了自编码器与传统降维方法(如PCA、FA、T-SNE、UMAP)的性能,结果表明自编码器在重建误差和均方根误差方面表现最佳,能更好保留数据结构,并提升模型泛化能力。该框架不仅减少人工干预,还显著提高数据处理自动化水平。未来随着深度学习与大数据技术的发展,自编码器结合生成对抗网络(GAN)或图神经网络(GNN)有望在复杂数据处理、实时数据分析和智能决策领域得到更广泛应用。

原文摘要 · Abstract (English)

This study proposes an automated data mining framework based on autoencoders and experimentally verifies its effectiveness in feature extraction and data dimensionality reduction. Through the encoding-decoding structure, the autoencoder can capture the data's potential characteristics and achieve noise reduction and anomaly detection, providing an efficient and stable solution for the data mining process. The experiment compared the performance of the autoencoder with traditional dimensionality reduction methods (such as PCA, FA, T-SNE, and UMAP). The results showed that the autoencoder performed best in terms of reconstruction error and root mean square error and could better retain data structure and enhance the generalization ability of the model. The autoencoder-based framework not only reduces manual intervention but also significantly improves the automation of data processing. In the future, with the advancement of deep learning and big data technology, the autoencoder method combined with a generative adversarial network (GAN) or graph neural network (GNN) is expected to be more widely used in the fields of complex data processing, real-time data analysis and intelligent decision-making.

自编码器降维自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。