An Introduction to Spark for Data Engineering | DataHour by Umesh Kumar
Skills:
Data Literacy80%
Key Takeaways
Apache Spark for Big Data processing, covering its high-level architecture, features, and advantages over traditional methods, with a focus on data engineering and analytics.
Full Transcript
foreign now before we dive into the spark architecture over about this part so first we need to understand what is the Big Data like why it came into existence like what was the need for the big data and what was the need for the spark we'll be discussing that so talk about the data the like before which uh start the history so first we need to know what is the Big Data again I'll be taking a very simple definition of Big Data like when we say big data one thing that comes to my mind is the huge amount of data which is in gigabyte terabytes and more than that to like it has basic three pillars variety volume and velocity when we say variety like the big data using the Big Data you can connect to a variety of data sources so that is my variety when we say volume volume is the large amount of data or the humongous data that it can handle or store then the velocity velocity means like it can process the data in a very small amount of time that is why these are the three pillars for the big data now what was the need for the big data so for that first of all we need to go to the uh like the computer so initially like during like initially what happened was during the industrial era when there was a lot of manual work so first thing was computer that came into existence for automation so if you think about the computer like the initial list so the two main operation that it does was first of all it was used for the basic data storage and second was the processing the task so during that data the primary means of storage was the files though like there were there can be any kind of file like it can be Excel it can be CSV or the text file or the word file so data could be stored in the files so after some time like when the data started growing so the companies wanted a central reposition that they have resistant for that relation database was implemented so relational database was better than storing the data in the first in terms so because it allowed data to be stored in the form of rows and columns that is in the structure format it also allowed the data to be stored in the normalized format such as to avoid the data redundancy and to form a land to remove the data duplicacy so this was working perfectly fine and these were most optimized for the read and write operations as well as the transaction processing however with the time like as the time started drawing data started also and it also started growing so at that time there was a need for one more solution like which could deal with the historical data and also could be utilized for the analytical and the reporting process for that data warehousing work uh they were optimal for the historical data and the analytical processing so this was working perfectly fine however with the Advent of Internet so data started growing exponentially why uh why because now a number of the user base for the internet was increased so people from all over the all over the country all over the world started using the internet that is why the data started growing exponentially so at that moment like the database and database these were the only solutions and it was difficult to store and process the large amount of data in a single machine because of the storage limitations
Original Description
"In this DataHour, Umesh will make you explore the world of Apache Spark and how it has revolutionized the traditional way of processing big data. Starting from the overview of Spark, including its high-level architecture and features, he will then explain the advantages of using spark over other technologies. After which, he will demonstrate how to use Spark for data processing with the help of a few coding examples.
For more amazing datahour session, visit: https://datahack.analyticsvidhya.com/contest/all/
Stay on top of your industry by interacting with us on our social channels:
Follow us on Instagram: https://www.instagram.com/analytics_vidhya/
Like us on Facebook: https://www.facebook.com/AnalyticsVidhya/
Follow us on Twitter: https://twitter.com/AnalyticsVidhya
Follow us on LinkedIn:https://www.linkedin.com/company/analytics-vidhya"
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Analytics Vidhya · Analytics Vidhya · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
The DataHour: Data Science in Retail
Analytics Vidhya
The DataHour: Anomaly detection using NLP and Predictive Modeling
Analytics Vidhya
The DataHour: Energy Data Science Project from Scratch
Analytics Vidhya
The DataHour: Explainable AI Need and Implementation
Analytics Vidhya
The DataHour: Google Cloud AI/ML
Analytics Vidhya
Prediction to Production in Machine Learning #machinelearning #prediction
Analytics Vidhya
Practical Applications of Data science in Ecommerce
Analytics Vidhya
How to tackle Overfitting?#machinelearning #overfitting
Analytics Vidhya
Building Data Pipelines on GCP #googlecloud #datapipelines #data
Analytics Vidhya
Hands-on with A/B Testing #abtesting #datascience
Analytics Vidhya
Efficient Implementations of Transformers #transformers #cnn #machinelearning
Analytics Vidhya
Modern Deep Learning Architecture #deeplearning #architecture #deeplearningtutorial
Analytics Vidhya
Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Analytics Vidhya
5 things you should know about Azure SQL #azure #sql #datahour #datascience
Analytics Vidhya
AI & ML in the Automotive Industry #machinelearning #ai
Analytics Vidhya
Building Machine Learning Models in BigQuery
Analytics Vidhya
NLP aspects in Telecommunication Industry
Analytics Vidhya
Practical Time Series Analysis
Analytics Vidhya
Fundamentals of Quantum Computing
Analytics Vidhya
A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
Analytics Vidhya
Classification Machine Learning Model from Scratch
Analytics Vidhya
Knowledge Graph Solutions using Neo4j
Analytics Vidhya
Model Guesstimation (MLOps)
Analytics Vidhya
ETL Pipelines in Google Cloud Platform
Analytics Vidhya
Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Analytics Vidhya
Getting Started with AWS EC2 #amazon #aws
Analytics Vidhya
How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
Analytics Vidhya
Certified AI & ML BlackBelt Plus Program #shorts
Analytics Vidhya
Visualizing Data using Python #machinelearning #visualization #python
Analytics Vidhya
DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
Analytics Vidhya
M in ML stands for Math & Magic
Analytics Vidhya
An Unsupervised ML approach using Clustering
Analytics Vidhya
Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Analytics Vidhya
Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Analytics Vidhya
Practical MLOps #mlops #datascience
Analytics Vidhya
Data Engineering with Databricks #dataengineering #databricks
Analytics Vidhya
Multi-Objective Optimisation
Analytics Vidhya
When Airflow Meets Kubernetes
Analytics Vidhya
AI in Banking
Analytics Vidhya
Learn Convolutional Neural Network for Image Recognition
Analytics Vidhya
Extracting Value from Data
Analytics Vidhya
How to measure Marketing Channel Effectiveness
Analytics Vidhya
Transforming Lives | Data Science Immersive Bootcamp
Analytics Vidhya
Stock Market Analysis - AI driven approach
Analytics Vidhya
Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Analytics Vidhya
Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Analytics Vidhya
The Power of Visualization | Tableau Full Course | Analytics Vidhya
Analytics Vidhya
Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Analytics Vidhya
Data Visualization in Data Science | DataHour | Analytics Vidhya
Analytics Vidhya
Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Solving any Machine Learning Problem | Approach and Steps Involved
Analytics Vidhya
Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Analytics Vidhya
Data Engineering in E-Commerce | The Best Case Study
Analytics Vidhya
Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Introduction to Federated Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Analytics Vidhya
Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Analytics Vidhya
Learn Hypothesis Testing | DataHour | Analytics Vidhya
Analytics Vidhya
A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
Analytics Vidhya
Making AI work for Business | DataHour | Analytics Vidhya
Analytics Vidhya
More on: Data Literacy
View skill →Related Reads
📰
📰
📰
📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Towards Data Science
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Dev.to AI
How Airflow is using AI to make data engineering more resilient, not more complex
Medium · AI
What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
Towards Data Science
🎓
Tutor Explanation
DeepCamp AI