对比机器学习与深度学习在本地和分布式环境处理高维数据的性能表现。
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
- 基于Epsilon、IMDb等数据集,测试不同模型在本地与分布式环境的表现。
- 使用Apache Spark在Linux上搭建集群,实现大规模数据的并行处理。
- 适合大数据分析与分布式系统开发人员参考实践流程。
本文报告了大数据课程中实施的一系列实践方法。工作流从处理Epsilon数据集开始,采用分组与个人策略;随后进行文本分析与分类(使用RestMex)以及电影特征分析(基于IMDb)。最后描述了在Linux环境下使用Scala实现的Apache Spark分布式计算集群的技术部署。
原文摘要 · Abstract (English)
This document reports the sequence of practices and methodologies implemented during the Big Data course. It details the workflow beginning with the processing of the Epsilon dataset through group and individual strategies, followed by text analysis and classification with RestMex and movie feature analysis with IMDb. Finally, it describes the technical implementation of a distributed computing cluster with Apache Spark on Linux using Scala.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。